Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2506.21876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21876 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:39.392445Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:04:43.145985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96d09218-983c-4fd2-8ab9-aa20774187c0 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:44.639025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.684245Z digest=sha256:576fcac536472d574ae7ca032d6145d5cfcbd3e77d3ac35126c55032328f94db

Observation 165c8bd4-30ce-460b-b4e2-37a3124c2fa5 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:44.365198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.707507Z digest=sha256:a2a66b727a736f6fe4fc941754fd70bfe8ac2be53551094fcb1cffb69d686f25

Observation 3b6f71db-95d1-49e5-9e86-1c708b75a125 · outbound

This paper cites The topological order of the graph is determined by time and world dynamics.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation The topological order of the graph is determined by time and world dynamics

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.089308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.734779Z digest=sha256:bce93fea399a784bf31bd14a422a4a75c4591800209e686098bf6c42a0c7c368

Observation afe4b3b5-5f30-46f6-b85c-2d8ea1bffb13 · outbound

This paper cites These multi-view tasks emphasize the model’s ability to synthesize distinct viewpoints into a coherent three-dimensional representation of object arrangements.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation These multi-view tasks emphasize the model’s ability to synthesize distinct viewpoints into a coherent three-dimensional representation of object arrangements

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.690342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.917251Z digest=sha256:7b788b0ef8ae2b69e7a1808480bad063519fb6b87f5ecf4aefd7492a1862004b

Observation 93f01c19-7bb6-40be-94fd-880cabfa4775 · outbound

This paper cites SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.200429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.200429Z digest=sha256:aa1d4b5731678dd015f582e1b4079a87086498169cbb4ba92f0723e3d8b1a12d

Observation 515bd20e-6045-42d8-9c10-e33ef8999192 · outbound

This paper cites Why think step by step? Reasoning emerges from the locality of experience.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Why think step by step? Reasoning emerges from the locality of experience

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.379328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.379328Z digest=sha256:b4cc847dde12f56b4051fcc4f2e1b0bba6ad561a636fbfd06ed64f54984bb83f

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:d6e3cc35d73f3f0563a49dc2a89026aeffd2ddfcbffc78ddd4edac9e763dc55a

Observation aeed72a1-4d00-48c2-8431-55b489ce3dc5 · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.457730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.457730Z digest=sha256:fd96f0aac61df56d4c0e2c05b170568e222a20d72347f342dfad379491ba7c38

Observation 8098a857-5480-4c86-831a-40b6fb1c9afb · outbound

This paper cites Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.511154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.511154Z digest=sha256:a42b47b563fb5c5cc3fc6b6306cf22e64a0cf38d42b7541bf6b6fa636d3bc835

Observation b42919ff-a75d-4e1c-bf9c-1db86f018aab · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.580685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.580685Z digest=sha256:2cd5aaf64b601de2e9783954a5f2f8aabc98be73f8456cb4dca7b0abb9447c3b

Observation 3ccfcbb2-0c3b-4b86-ac09-df84f94e9a8b · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.649660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.649660Z digest=sha256:528544fd9a809165a77462c3046c10d6cce5e6bef22e8c3cd9985957e9031f28

Observation d2eb0d2c-2e7e-46ed-b259-451264b35f48 · outbound

This paper cites This task evaluates whether the model can accurately discern spatial relationships based on visual cues.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This task evaluates whether the model can accurately discern spatial relationships based on visual cues

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.479831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.794107Z digest=sha256:3a628543cce741d1ae6bd905187b5e7fa9a3d0ed57fb566b85b51155cc3fb77c

Observation f3fab12a-a7f6-4d52-bcfc-09507104ee3b · outbound

This paper cites Thus testing the model’s understanding of spatial constraints.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Thus testing the model’s understanding of spatial constraints

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.162217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.824097Z digest=sha256:08202cca328a5c746fb3f4aed3e767845dd3c95b1c7ce87825ee235ae4ea5384

Observation e601943c-4656-44f8-99b2-02a51961ab65 · outbound

This paper cites Thus testing the model’s understanding of object size.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Thus testing the model’s understanding of object size

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.943046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.864031Z digest=sha256:324f8569bce6497dd4adf94dc171c998eb45110adec091cd80886da0eac3e9e1

Observation 155d7142-72dd-44bc-8a7c-1ce312aabd2d · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.551464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.953529Z digest=sha256:997e781ed1dd833a0b96f48aac679dac31eef4a89c9da7efb73359d9c397eac4

Observation 47f848a7-016a-4c1b-9e7f-817bd4eb82cd · outbound

This paper cites Collectively, these tasks investigate the aptitude of a model to maintain consistent temporal representations, estimate durations, and infer the correct order of events.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Collectively, these tasks investigate the aptitude of a model to maintain consistent temporal representations, estimate durations, and infer the correct order of events

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.369495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.991794Z digest=sha256:23f604c67a124e156dcb4951958ac34e791c33e2002deb5db739d55ff988c4f0

Observation 2a7c8717-0104-4d8b-b9ce-b329c2b846e8 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.257805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.024311Z digest=sha256:b347471d848750d62199bb023182f0170e11fbd35cc1ebdefbd6eb760f7b4206

Observation e521baee-d6a4-4daa-b710-1e437ba9f704 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.135747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.094325Z digest=sha256:9b24f08928fe721795bd9d1f4f74fa32cfff3accc6bb6587903aa83d89055612

Observation 2d1df94b-2d08-4c4c-bc53-87f17ad2d684 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.036339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.154224Z digest=sha256:fff84facb77bd92990b7fa7b0688867d208a2e5f75f4393344dfe84442001729

Observation 42c72287-8263-46a2-b8fb-3256ec1db670 · outbound

This paper cites This setup tests the model’s capacity to identify the moving action of objects.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s capacity to identify the moving action of objects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.857927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.224945Z digest=sha256:b50e6793e24a1f041721e64c14e9e7964ba5df7d2871441a91596a3a1e1141ad

Observation e8c37cfd-8fe3-4aa4-a7dc-816b6751e3c5 · outbound

This paper cites This setup tests the model’s capacity to track position changes over time and estimate relative velocity.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s capacity to track position changes over time and estimate relative velocity

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.741165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.287916Z digest=sha256:65e4ba613f75993017581d1f9eae5aae45faade3eb73b01383ea96da9242833f

Observation 958617c8-47ad-4c21-ae53-9e26396b54d7 · outbound

This paper cites This setup tests the model’s 20 capacity to track position changes over time and estimate relative moving direction.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s 20 capacity to track position changes over time and estimate relative moving direction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.615188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.367788Z digest=sha256:fc79dcfbd7a6440b58407d394db520a4c48ad235d177565a90ff789d0861c7e2

Observation 6c57eddc-4f41-46e4-b015-a3257e87c33d · outbound

This paper cites Quantitative Perception.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Quantitative Perception

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.495654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.459906Z digest=sha256:45a33a536baafd30df201143dbcdb47f591c24be4e7146a6b759f7f79249025d

Observation 8b0665ea-32c4-4eb6-b3a6-141e4b19dccb · outbound

This paper cites This setup evaluates the model capacity for discrete numerical estimation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity for discrete numerical estimation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.375521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.534512Z digest=sha256:757df310f6a08d5acc3ffafee66fa7b70acc58f1c773a57444bf899dd9d1c0dc

Observation 74281348-1b85-4e0e-aee5-7c50ec805ee0 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:41.268379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.610245Z digest=sha256:0593c85e9937c75cc390d3cadfd8cbdfe14b0e09db185bef450ec10443f71986

Observation e73263ab-5223-4843-a247-4e8fe27ccb9e · outbound

This paper cites This setup probes counting skills, numerical reasoning, and perceptual comparisons in a visual context.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup probes counting skills, numerical reasoning, and perceptual comparisons in a visual context

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.156533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.684429Z digest=sha256:eff3b53538d71c635be0bc0f3bbfca46ae364f8e5de18dae40b90a9d6795dc46

Observation 5f530ea3-642d-45e2-a267-7e1b50930ba5 · outbound

This paper cites This setup evaluates the model capacity in physical reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in physical reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.056090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.774412Z digest=sha256:0f1b07c80256e6e5b651d091eb5d0e9b25b5e6ebd25f2f135c16802e15c808fc

Observation 4ffb2d8c-db14-4262-8d1a-663616202112 · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.842243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.869516Z digest=sha256:8dd4b3a40f2641135c251964cc28114a9f41bde93d3302b792b2d1a7decb22ca

Observation f77b19c7-0baa-4c21-93b7-78fb3eb07adc · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning for robot manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning for robot manipulation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.672138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:38.977727Z digest=sha256:ed4873b4790c4f65c7ba9b13e61dedbbc4782d450176183cf6bf546f75b9b074

Observation 241f8ca8-00e2-4180-a1c0-d8ec06f2aceb · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning for autonomous navigation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning for autonomous navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.525701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:39.048655Z digest=sha256:00735db7e7455a7d9a049e52499614b0af56288a8fe82392fa01214a56ee602f

Observation 0ee026b0-4639-4502-97dd-395333f84139 · outbound

This paper cites This setup evaluates the model capacity in multi-step predictive reasoning for robotic manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in multi-step predictive reasoning for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.350422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:39.122549Z digest=sha256:d153e1efc4ba42c582a5a67c5a56ec01be9bbca082f84f347aafe744888dd4fb

Observation e656d49b-4ad5-447f-b6f2-a0f0e3964994 · outbound

This paper cites This task evaluates the model’s ability to perform compositional inferences about physical causality and object behavior.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This task evaluates the model’s ability to perform compositional inferences about physical causality and object behavior

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.162338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:39.211208Z digest=sha256:3d480e6fc7fbf1df8ea0ef0a1f05facf426ee2202941ea2f154d397552e14e05

Observation 6de38852-8989-4b01-8495-8a41ab86f89d · outbound

This paper cites This setup evaluates the model capacity in concurrent action predictive reasoning for robotic manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in concurrent action predictive reasoning for robotic manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:39.987746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:39.298648Z digest=sha256:5a29c5b33194f6e0aab28db83dcc26cd90d37e9c60ef425414cbe3435d15983f

Observation dc74911f-a30f-40e4-8d70-b6e3e09c7389 · outbound

This paper cites yes” or “no.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation yes” or “no

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:39.809002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:39.392445Z digest=sha256:2451fa1d3d993f20fed24b987fe0679bc95be74112ec427342f396280e6635d4

Observation af3719b3-b7c1-466c-8f02-660250ebc0c0 · outbound

This paper cites John Wiley & Sons Hoboken, NJ.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation John Wiley & Sons Hoboken, NJ

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.324343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:36.849611Z digest=sha256:c1f32c7feac81d9b00db97a8cb69ee4ae7ff92f556faea3ed16c8444ff55d883

Observation cab77872-0d73-4b5b-b584-3375f1655071 · outbound

This paper cites S(n) t to denote the whole relationship between n component states and complex states.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation S(n) t to denote the whole relationship between n component states and complex states

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.816986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.758890Z digest=sha256:eac9794af0778fdd61dc6def7fff0bb92bb97a78d17f368b255e199d80210a0c

Observation 399c084c-8c05-4a5e-b08e-811189840f70 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation GAIA-1: A Generative World Model for Autonomous Driving

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.281149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.281149Z digest=sha256:2633e0439d64186bdecae8ba83e89ad9d497b490e9aebc51d54041066c21a1c2

Observation ee900967-a049-4644-9de2-b0cda7747bd3 · outbound

This paper cites Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.339688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.339688Z digest=sha256:20333fca353170727c0dc34115dc2328628ab70fad74475305ab985a5923efac

Observation 2f2ff74f-0605-4c25-919e-08b5cca1b525 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.559873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.559873Z digest=sha256:5734d1d74651c23a1b16a35c908b9833bfa3e0affe601359bcb21da14f87ec2a

Observation 7f3f7446-c88c-416c-a422-478b4e23c352 · outbound

This paper cites SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition

Reference 2019

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:22:39.582059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.555217Z digest=sha256:5c43de6571ea8a2532b8776cc9e4192ecc3be3b063f72bded0a0bb6bf6fb16a0

Observation 890028db-ee63-46c2-ab8e-277326e274e3 · outbound

This paper cites In Advances in Neural Information Processing Systems, volume 33, pages 10514–10525.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation In Advances in Neural Information Processing Systems, volume 33, pages 10514–10525

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:02.907075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.101582Z digest=sha256:41334e6b3cd966d06e76df3588235025acc7eb9f136436e32e01639b3850ef5e

Observation 6c279471-d033-4ad8-8d46-16a9da93dcc7 · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.996473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.996473Z digest=sha256:826755c22c3cd1df41f223bdb19811eb04809a6f1c33c4a4f65e624f4b0bc5b7

Observation bc149e21-7257-48a7-ad3a-38db00cca78b · outbound

This paper cites Faith and Fate: Limits of Transformers on Compositionality.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Faith and Fate: Limits of Transformers on Compositionality

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.904449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.904449Z digest=sha256:0f43f32882cfe20d8d80965074b950efba6c927e9833b88c4d1405012a77c5d0

Observation 802dd24f-d6f8-43b7-b55e-d9e7c290a502 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation How Far is Video Generation from World Model: A Physical Law Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.307022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.307022Z digest=sha256:d4739773cd3dce02e845e19592ec078b1350fae230309adfe716326e4cff0628

Observation 3392a115-99fe-46c2-9cc2-5a7dc56cf3e8 · outbound

This paper cites In The Thirteenth International Conference on Learning Representations.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation In The Thirteenth International Conference on Learning Representations

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.942959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T22:22:37.626475Z digest=sha256:123d286a60247898407f7ade14724984fb984c2cf15b53dae69608811a71850b

Pith citing papers

Observation 770e4f03-85ca-428b-9829-82fbbaaf69ce · inbound

Egocentric Bias in Vision-Language Models cites this paper.

Egocentric Bias in Vision-Language Models Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.145985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.145985Z digest=sha256:d46b1cc9db35a060ea6c5946c75beca7ec1cd5c3b1ea6ccee30c4e019307bb53

Observation 17d4b390-cd08-4e15-a2f8-2390ef800198 · inbound

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics cites this paper.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T21:36:42.378149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T21:36:42.378149Z digest=sha256:1685a213070dee003a055a85646df5068e3931536d60b1c24d1333b0b93f7586

Observation f6bf9e4e-5780-48d2-9252-45f48f43f995 · inbound

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning cites this paper.

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T07:27:08.281893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:27:08.281893Z digest=sha256:17146972ea296c9aa3b3bf79327982664362ae1ae8862a71db80a6734072a4cf

Observation d023fc94-0bd3-4f98-b89b-fe4b2cc8c1be · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:36.966727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:36.966727Z digest=sha256:e2fd6c577cbd4959562e93e43ff24ddcfc4d6b416038eb449fdaa9cb9d48586e