Pith. sign in

Paper Citation Record · LEDGER

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.06006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06006 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:07:24.277177Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8863294b-c256-4205-9463-1d4538c03e41 · outbound

This paper cites Recurrent world models facilitate policy evolution.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Recurrent world models facilitate policy evolution

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.201072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:23.988932Z digest=sha256:9aa7aa5987c5fced0f4a78b90db04cc853a66e28f922f512859e675e51cc7dc1

Observation 4286e6d7-37ae-4216-b2cc-2fbec26dba99 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Cosmos World Foundation Model Platform for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:23.995257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:23.995257Z digest=sha256:e3f8e89024942895ef0827e85b7b7ef8c56a3eea01e060b91144c9d06308a38a

Observation 0ca95a54-c6cc-4406-a062-1c01f424057b · outbound

This paper cites Genie: Generative interactive environments.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Genie: Generative interactive environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.001012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.001012Z digest=sha256:52e7650a64e923ecda8986eeb39b6f4dd9cd1ad8b09dc1c55a0247ff9f7297fa

Observation 7d0c545d-536c-4143-ba7b-cbf786b12461 · outbound

This paper cites Video generation models as world simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video generation models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.173532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.006263Z digest=sha256:c1e3c52be5f661da29158facfeee8d474ac96ab39e6e574f17d284b5d0b9aba1

Observation ea51fd94-ebf7-46bc-bed2-42afdc6a2250 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics WorldSimBench: Towards Video Generation Models as World Simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.010922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.010922Z digest=sha256:b0319377b268b92819f09155a2f74f5c0a1e96fdf5e31856d7fa484333ad19ac

Observation 7c89477f-64fd-4e9c-acd9-ef939b47b253 · outbound

This paper cites Do as I can, not as I say: Grounding language in robotic affordances.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Do as I can, not as I say: Grounding language in robotic affordances

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.155966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.015924Z digest=sha256:ab7a0ac699882b64f3e006f94588bf971fae68af8eaba52c030aab3ffa1a3ce1

Observation 27e59029-a26d-4f79-bcae-5794c0d4232c · outbound

This paper cites Inner Monologue : Embodied reasoning through planning with language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Inner Monologue : Embodied reasoning through planning with language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.137732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.020456Z digest=sha256:27ceb7c7a95ccf9e19b358cc489e197c77e325827a05363066733895278274d7

Observation 3092e8bf-2fe7-43ca-b59b-c0bbd45e614e · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.025967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.025967Z digest=sha256:e30e252b842fe62452cc205a261e800ea2a6e25d6d4acceea099d3f2cda0c276

Observation 2a29bc56-079a-476f-adc6-e35ffee01026 · outbound

This paper cites A Generalist Agent.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics A Generalist Agent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.031045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.031045Z digest=sha256:40917edc213ce36167e0f7aff1da4bce1608fd50e3493e5069e55fce6600e6d3

Observation 32537f05-b388-4691-98de-c7c20ef1cf32 · outbound

This paper cites Learning Interactive Real-World Simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning Interactive Real-World Simulators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.035996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.035996Z digest=sha256:e9077df7395900fe2ac501542e05e62fbd21581ea19ed819ebf1fac3d1f7ea2f

Observation 78cea173-cdcd-4395-8461-ecbc2d080cdc · outbound

This paper cites Mastering diverse control tasks through world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Mastering diverse control tasks through world models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.122718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.042286Z digest=sha256:1ae756be707602cd1b607d3e65930a24434d044c0cceea8bcef76ebeab837af3

Observation d0ab5390-6057-4e67-ae4b-ba144b5846da · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.046755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.046755Z digest=sha256:f4430449433feda62792885e1925803cbe8de601949ebbbc09ebc51670a80cd5

Observation d3c9dccb-f84f-44ff-a2fa-27e1dc9191db · outbound

This paper cites Do generative video models understand physical principles?.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Do generative video models understand physical principles?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.051408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.051408Z digest=sha256:e8140fb592d6d428691c30390483a1cf81acf33860506415edab382c15e3804b

Observation 6cf3a3ba-b5db-450e-806f-021a7abf32c7 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Physically grounded vision-language models for robotic manipulation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.107730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.056924Z digest=sha256:cbd1539bc40439714e3acc1aa181811d20619281fc73d5dde3cbea8ebf2b9ea4

Observation 69c4f1b3-c04f-4d8c-9e03-c6280b7f88ff · outbound

This paper cites an unresolved cited work.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:25.091676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.061561Z digest=sha256:151b640f624682c80306dd37fe7cc722ab3eba37a2e03b1b802c8d22dd7d470a

Observation 85becdf0-b936-4438-b842-bbc53d393128 · outbound

This paper cites Can language models encode perceptual structure without grounding? A case study in color.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Can language models encode perceptual structure without grounding? A case study in color

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.075935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.066153Z digest=sha256:f8c581b2409208a75f05edfa72abc0a4693228feeea70af9febe2e01a43b9b46

Observation 3a0aab83-bf9d-49ef-9bf7-5f8aecfbd6df · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.072246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.072246Z digest=sha256:c886e7da02b76e1f65bfb44656b5a51ed370c9a32e18ade163864d1dfb060ca1

Observation 0f58401b-5e6a-42a0-aa18-07a5686137cb · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.060218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.078717Z digest=sha256:13840794d663cc696f7587914432d797a79d7c94cd11c6d8d3e5e6a00aa260a0

Observation 01bca877-3c16-4f23-92ba-7269b995fbf4 · outbound

This paper cites Video PreTraining (VPT) : Learning to act by watching unlabeled online videos.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video PreTraining (VPT) : Learning to act by watching unlabeled online videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.043321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.084572Z digest=sha256:4ce0b14d654668e72190278e3c9e80af0e59c20f3762d832ad4a85dd95f0611b

Observation 59a6acba-88a1-4f93-8150-cd732c3975b0 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Moments in time dataset: one million videos for event understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.027199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.094006Z digest=sha256:d6d69ea1ad45e9b87924f63fbe58761fd95770131ced1787c42c3af8943f7e65

Observation 6f02e7ac-5174-4f45-ab1a-a94afb616404 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics The Kinetics Human Action Video Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.106289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.106289Z digest=sha256:e87ec7a07789cde377e0ab719b653f56245543e3e89fa3d6deaa40c5ba876284

Observation e580f276-799e-4ede-9b90-7b96d4523b84 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics A Short Note on the Kinetics-700 Human Action Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.112945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.112945Z digest=sha256:ea3b5d859171474ea10f29f4a876865ada5e930f4b2431e5485feafacae639bd

Observation addd9681-e6dd-4c3d-9596-6b0341cff92a · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.117780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.117780Z digest=sha256:e6bbc33d644ae9afa88a223927da0685059ce459287ec35a7a11834f5800a0c1

Observation 845beb56-b69f-4cf9-b273-296b9a23170e · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.123574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.123574Z digest=sha256:926e860ee4884ac9efce56b18a99f22675583bc04fd7f582b9a27367526a780b

Observation b7a16e98-9bca-4d02-9e3f-2ebaeb863cd8 · outbound

This paper cites Scaling egocentric vision: The EPIC-KITCHENS dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Scaling egocentric vision: The EPIC-KITCHENS dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.009280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.131233Z digest=sha256:dba033375282eabe938f637c4bb3768df1339a4322320c4fc86356e90e0aaab8

Observation 8b21755f-24ef-46e9-81fb-a454489583d8 · outbound

This paper cites s1: Simple test-time scaling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics s1: Simple test-time scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.139015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.139015Z digest=sha256:7cfb8b1e67ab52100c3e7cf410d10f4e14ce7b590f76561dc853a5c247963711

Observation cee2bfe3-7a20-43d3-9550-c4581fa7b974 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.146649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.146649Z digest=sha256:cc1a2171aa9cd0bd310de42a50f6af54eac7a0346bfae796a61ae54b01499a8b

Observation abc8ad40-06ad-4a4f-9ead-752ccaf4dcea · outbound

This paper cites Kubric: A scalable dataset generator.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Kubric: A scalable dataset generator

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.993415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.154506Z digest=sha256:02614dd3c546a02efaf8a6203d001fa611e56ad475405b6949bfa16b8ed98ef9

Observation d90e9194-b20e-49c3-aa98-da96adc23f23 · outbound

This paper cites InstructPix2Pix : Learning to follow image editing instructions.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics InstructPix2Pix : Learning to follow image editing instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.978148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.160136Z digest=sha256:aea7ed242709142b91e2a8e32a49cfb496afde16787d339da62de352159536cd

Observation e85fb9d6-0103-4439-b9cc-e33646806b97 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.165875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.165875Z digest=sha256:36460306094b6009566f950d7897e9dec117f8ff47585401a3bf564b8465ac56

Observation e0c0d67e-b944-490e-99fc-7cb50afc0217 · outbound

This paper cites SmartEdit : Exploring complex instruction-based image editing with multimodal large language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics SmartEdit : Exploring complex instruction-based image editing with multimodal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.961540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.171421Z digest=sha256:fdb78a6ae4a5816755d881ecfb0526333d1d02c84438d7b64f171669bd02eca1

Observation 9f064c14-cceb-4b80-a95a-21b137f4c6f1 · outbound

This paper cites BERTScore : Evaluating text generation with bert.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics BERTScore : Evaluating text generation with bert

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.943531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.177171Z digest=sha256:5dce608a260502887d1b5eb1a91dcbd4bdbfb1b2281fc9eeaece01dbab4d6388

Observation 3a054d53-0926-4a4b-8846-47f5df5cce80 · outbound

This paper cites ROUGE : A package for automatic evaluation of summaries.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ROUGE : A package for automatic evaluation of summaries

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.926323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.181794Z digest=sha256:040cf266f43457a56ce82d51f2cebe867b02ab7b1f5c3993456fae055dae47fb

Observation 40bd7703-77e1-4a64-8af0-24e30a93a50c · outbound

This paper cites BLEU : a method for automatic evaluation of machine translation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics BLEU : a method for automatic evaluation of machine translation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.910375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.186374Z digest=sha256:d052a7e7e9d3dfaf16c4b185fdab98c1d315e45e7f4c9299a1606d2549f9c0c4

Observation 22b337b0-2348-4812-b120-9203aac9fc09 · outbound

This paper cites Variational best-of-n alignment.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Variational best-of-n alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.894651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.192078Z digest=sha256:d202aac0ab8a5bba2732d9d8fbe948ca69a35b2a3397e7bee2db949b0c6367d6

Observation 04175720-5e5d-4a3d-bbb1-8affe0c9027f · outbound

This paper cites Learning to predict by the methods of temporal differences.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning to predict by the methods of temporal differences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.879924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.196626Z digest=sha256:cb43c7bcd4236f6cb3d2997601903bf066b7ace8fb8817c9019eb98a76d90570

Observation d3207217-fac6-436e-a918-954b2067100b · outbound

This paper cites Learning latent dynamics for planning from pixels.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning latent dynamics for planning from pixels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.863771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.201635Z digest=sha256:330c3520bb1d8d7ddbc8fde82d8d6482a7296c8247388f39f623c6fc69b6db1e

Observation e934c709-0463-4043-b79d-342cc8da7aa4 · outbound

This paper cites Transformers are sample-efficient world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformers are sample-efficient world models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.848709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.206106Z digest=sha256:5cfc6f86c5c0f7c180b907805f8748e06afca3a56467bb5868640d7280f8a895

Observation f5d5c6b4-4972-4c15-90a1-8e701d569918 · outbound

This paper cites Transformer-based World Models Are Happy With 100k Interactions.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformer-based World Models Are Happy With 100k Interactions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.210816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.210816Z digest=sha256:fe96d421e0fbd4a1d17235ca5aaa9a7fe78cf53a910d6371a928cc00b829b562

Observation 92f9b374-e22a-4e4f-8497-f805085aa920 · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Diffusion for world modeling: Visual details matter in atari

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.832798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.215773Z digest=sha256:dd6c8350dd943aaec7ed2064633eaef6a5efb67b6436b55e3e5abe9eecfcc43b

Observation 1ec03a46-1e03-43c3-9ea4-a5a150fe8928 · outbound

This paper cites Video-LLaVA : Learning united visual representation by alignment before projection.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video-LLaVA : Learning united visual representation by alignment before projection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.816123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.220051Z digest=sha256:c1674a5229f6cc91a3d8b9ee47eb9fd832907bd50c46314b2b72bb5674b77815

Observation 0ee50795-2a17-4ce8-b0d3-d58b82240769 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.224463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.224463Z digest=sha256:72a1df152032bdb4a00c91eb22e48471dc6f6e24c4a56ea37778909652206649

Observation 705237c3-ef8a-4da8-ab19-11f7323d799e · outbound

This paper cites iVideoGPT : Interactive VideoGPTs are scalable world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics iVideoGPT : Interactive VideoGPTs are scalable world models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.799182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.229666Z digest=sha256:b817b188f765fd09a8df892283d8defd3b916d66ac41495993a2ca516d576a37

Observation 3d3f5554-3e1e-4183-b63f-9c8d3a053a5e · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Decision transformer: Reinforcement learning via sequence modeling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.780990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.234370Z digest=sha256:aaa0f7366f0d4ee2693cf205ef7473bff383a55f6b8752a6eb6294dedf38f871

Observation 0d753bfb-16d2-407e-bd88-cf44f27b7454 · outbound

This paper cites Vision-language models provide promptable representations for reinforcement learning.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Vision-language models provide promptable representations for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.764728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.238717Z digest=sha256:834903e51feb6256bf94a8d28d355b14fd039e2866089ff2281bae37ed994e36

Observation d8e610c3-5eba-4198-89e6-b50ebdb6ae98 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.748169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.244183Z digest=sha256:149032426d7d39922d67cdedc82330521bd56a620589828933b86f45d7d327cd

Observation bd6a97b2-880f-4f4a-81a5-1f4dba080336 · outbound

This paper cites Video as the New Language for Real-World Decision Making.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video as the New Language for Real-World Decision Making

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.250329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.250329Z digest=sha256:26e6e5b80243d8d4e44206ba3c75e8550dd8942e9a3dd4b2cb29963733ae71c4

Observation 13d75b34-ae20-4441-806d-4682344e22b4 · outbound

This paper cites VideoAgent: Self-Improving Video Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics VideoAgent: Self-Improving Video Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.256971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.256971Z digest=sha256:9ff137ebaa264ab5641d76a8c2127b3cbfcdc718d478868dd5873543233843da

Observation e3caf459-b8b1-403e-a755-12a927bd2ff1 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.262152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.262152Z digest=sha256:174fc438deb429a58f00be50eec4a6845852c6e14aa5481e0d7af83b4c75f9e5

Observation 32172cfe-5218-4da7-9b8b-0b6637d226e4 · outbound

This paper cites LoRA : Low-rank adaptation of large language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics LoRA : Low-rank adaptation of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.731994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.267271Z digest=sha256:a84c1bdcae6641540e4d0b473ebe76879ab4dd579847f5c6d4206314c57c93a4

Observation 28e71d01-3eb4-46c8-9e6f-0f23089554fc · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformers: State-of-the-art natural language processing

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.715787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.272518Z digest=sha256:78acfc7ea3cad77d4123992250e335fcddaec75762cfb3a8dce4eab2a1e620e3

Observation 24fe666c-1b17-4bb7-8922-5bf0e22de8bf · outbound

This paper cites write newline.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.277177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.277177Z digest=sha256:4be26b78fd723792cf2aeb18811252754e4e0fc24819ee2a2f51c1b50ee3087c

Pith citing papers

No inbound Pith citation observations are available.