Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2506.21876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21876 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:39.392445Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:04:43.145985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96d09218-983c-4fd2-8ab9-aa20774187c0 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:44.639025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.684245Z digest=sha256:23a938f48ca005ccb510998705f0dbefb9abaab6cf8978c962cdfc7c245b24cd

Observation 165c8bd4-30ce-460b-b4e2-37a3124c2fa5 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:44.365198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.707507Z digest=sha256:1affd3e2d9583b6bf231cd52cc7c5253bb684654a6b5e399957aed68937f9492

Observation 3b6f71db-95d1-49e5-9e86-1c708b75a125 · outbound

This paper cites The topological order of the graph is determined by time and world dynamics.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation The topological order of the graph is determined by time and world dynamics

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.089308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.734779Z digest=sha256:5cdf6d623f29d09403f16855c93fdbfd77f820420baff29319d0bda53e08bac9

Observation afe4b3b5-5f30-46f6-b85c-2d8ea1bffb13 · outbound

This paper cites These multi-view tasks emphasize the model’s ability to synthesize distinct viewpoints into a coherent three-dimensional representation of object arrangements.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation These multi-view tasks emphasize the model’s ability to synthesize distinct viewpoints into a coherent three-dimensional representation of object arrangements

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.690342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.917251Z digest=sha256:ab7acfdc4b1c755804188b9f56fb518b4110abb0022b88410a85d7be5cecdc13

Observation 93f01c19-7bb6-40be-94fd-880cabfa4775 · outbound

This paper cites SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.200429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.200429Z digest=sha256:f291f49a1d8186965be10ecf129cd5e05964018f15dde827b7d91e41a8af61d8

Observation 515bd20e-6045-42d8-9c10-e33ef8999192 · outbound

This paper cites Why think step by step? Reasoning emerges from the locality of experience.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Why think step by step? Reasoning emerges from the locality of experience

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.379328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.379328Z digest=sha256:2a2b90906262d1fee8a4479ab9470f52cfaea7a51de6e098217f3202bd9cdab8

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:2c3306fd691200d884e97bd41c8a73e8765c58ed8c607148812c792a26f06016

Observation aeed72a1-4d00-48c2-8431-55b489ce3dc5 · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.457730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.457730Z digest=sha256:fd96f0aac61df56d4c0e2c05b170568e222a20d72347f342dfad379491ba7c38

Observation 8098a857-5480-4c86-831a-40b6fb1c9afb · outbound

This paper cites Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.511154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.511154Z digest=sha256:005494c367ad1e63bbb23e34795152d74bc6d32bcc73ac3a872726329bf954b7

Observation b42919ff-a75d-4e1c-bf9c-1db86f018aab · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.580685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.580685Z digest=sha256:2cd5aaf64b601de2e9783954a5f2f8aabc98be73f8456cb4dca7b0abb9447c3b

Observation 3ccfcbb2-0c3b-4b86-ac09-df84f94e9a8b · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.649660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.649660Z digest=sha256:528544fd9a809165a77462c3046c10d6cce5e6bef22e8c3cd9985957e9031f28

Observation d2eb0d2c-2e7e-46ed-b259-451264b35f48 · outbound

This paper cites This task evaluates whether the model can accurately discern spatial relationships based on visual cues.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This task evaluates whether the model can accurately discern spatial relationships based on visual cues

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.479831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.794107Z digest=sha256:3492cec6b5bcd3602e126f95b54d6a4db00dcd43006719c40475bcd3c4c583ac

Observation f3fab12a-a7f6-4d52-bcfc-09507104ee3b · outbound

This paper cites Thus testing the model’s understanding of spatial constraints.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Thus testing the model’s understanding of spatial constraints

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.162217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.824097Z digest=sha256:23cbee9b1f13f8af633e7e1c0043242f33dd9fe5a7875fa54d29a048438e3de0

Observation e601943c-4656-44f8-99b2-02a51961ab65 · outbound

This paper cites Thus testing the model’s understanding of object size.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Thus testing the model’s understanding of object size

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.943046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.864031Z digest=sha256:4b38d73b8ffc6aa65c9717cff281539e5e55d020e55286aec1520f5cf7b3db2b

Observation 155d7142-72dd-44bc-8a7c-1ce312aabd2d · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.551464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.953529Z digest=sha256:cf51d0a07a6e6a203ed4f1e4afec98ac511b8364c92b98651c6e22b21909138f

Observation 47f848a7-016a-4c1b-9e7f-817bd4eb82cd · outbound

This paper cites Collectively, these tasks investigate the aptitude of a model to maintain consistent temporal representations, estimate durations, and infer the correct order of events.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Collectively, these tasks investigate the aptitude of a model to maintain consistent temporal representations, estimate durations, and infer the correct order of events

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.369495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.991794Z digest=sha256:0f0d68b25c8320e85ba191f246596db0ceed6ef254f43e1aba9577c5ab7aa8e5

Observation 2a7c8717-0104-4d8b-b9ce-b329c2b846e8 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.257805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.024311Z digest=sha256:3286abd1aaee307075f879a7689cdc5066ce8a7ca4fb0fda18e47699d20dec37

Observation e521baee-d6a4-4daa-b710-1e437ba9f704 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.135747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.094325Z digest=sha256:2a9cdd06bd3c8c47fd33a5b01440a046ab9325a208291637a4e4d6508c901acb

Observation 2d1df94b-2d08-4c4c-bc53-87f17ad2d684 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.036339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.154224Z digest=sha256:6c22ae9cc9da7193f3d82832b5036ef4f080834d40674e8f18369e41572c6173

Observation 42c72287-8263-46a2-b8fb-3256ec1db670 · outbound

This paper cites This setup tests the model’s capacity to identify the moving action of objects.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s capacity to identify the moving action of objects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.857927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.224945Z digest=sha256:6dc6c151d4371471aa4a2e223e7d0b67b675ecd1ddbe2697c4adcbdbe4f32a94

Observation e8c37cfd-8fe3-4aa4-a7dc-816b6751e3c5 · outbound

This paper cites This setup tests the model’s capacity to track position changes over time and estimate relative velocity.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s capacity to track position changes over time and estimate relative velocity

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.741165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.287916Z digest=sha256:7e27b6689d7d28b3cd768e70c03a44b29e91d4fdec496df81fa20cbd88ae53a4

Observation 958617c8-47ad-4c21-ae53-9e26396b54d7 · outbound

This paper cites This setup tests the model’s 20 capacity to track position changes over time and estimate relative moving direction.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s 20 capacity to track position changes over time and estimate relative moving direction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.615188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.367788Z digest=sha256:ddcb13982327c54ac33d93d53c8a107011b5b90476675d933d40d3e445936ec2

Observation 6c57eddc-4f41-46e4-b015-a3257e87c33d · outbound

This paper cites Quantitative Perception.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Quantitative Perception

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.495654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.459906Z digest=sha256:995c8723254b5cb04a31769ef3cf009d48598d9776d0080500327b9800037917

Observation 8b0665ea-32c4-4eb6-b3a6-141e4b19dccb · outbound

This paper cites This setup evaluates the model capacity for discrete numerical estimation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity for discrete numerical estimation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.375521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.534512Z digest=sha256:f4c7baaaa9b704d98aa37b236c442c6e0b7b2911282c65e0b850890c954b0545

Observation 74281348-1b85-4e0e-aee5-7c50ec805ee0 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:41.268379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.610245Z digest=sha256:1c31821e54582ec98f48a4bd6cc6ccb12a2776161fc0230673c208f6da067d53

Observation e73263ab-5223-4843-a247-4e8fe27ccb9e · outbound

This paper cites This setup probes counting skills, numerical reasoning, and perceptual comparisons in a visual context.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup probes counting skills, numerical reasoning, and perceptual comparisons in a visual context

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.156533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.684429Z digest=sha256:c12b6751316aca4a0a4facfe74ca6bfa5fed9dd41d92be41accd9feba3bc75e5

Observation 5f530ea3-642d-45e2-a267-7e1b50930ba5 · outbound

This paper cites This setup evaluates the model capacity in physical reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in physical reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.056090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.774412Z digest=sha256:d0a0effa687092b275b133139ca9ab0d0d27cc513c5f748146d200ca3d6bfa62

Observation 4ffb2d8c-db14-4262-8d1a-663616202112 · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.842243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.869516Z digest=sha256:01ac1216cd123d0f53cb021fe6ad9c4ae0ff12ae574b1d7c2490908ad6fa93c5

Observation f77b19c7-0baa-4c21-93b7-78fb3eb07adc · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning for robot manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning for robot manipulation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.672138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:38.977727Z digest=sha256:ad9fc15784ff7a575d0b0d7c7cd49d08e2e7eb042c8759a606390867290070f3

Observation 241f8ca8-00e2-4180-a1c0-d8ec06f2aceb · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning for autonomous navigation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning for autonomous navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.525701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:39.048655Z digest=sha256:00bdbe2107e55103e77efa35787e299596277b5571300257829ae672e3e8b277

Observation 0ee026b0-4639-4502-97dd-395333f84139 · outbound

This paper cites This setup evaluates the model capacity in multi-step predictive reasoning for robotic manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in multi-step predictive reasoning for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.350422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:39.122549Z digest=sha256:649fb5d1bb1600c6f9adf5dae94a243783630beba4446f728833dcf76d5c6f4f

Observation e656d49b-4ad5-447f-b6f2-a0f0e3964994 · outbound

This paper cites This task evaluates the model’s ability to perform compositional inferences about physical causality and object behavior.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This task evaluates the model’s ability to perform compositional inferences about physical causality and object behavior

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.162338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:39.211208Z digest=sha256:2313876c14450eafdbfe57a3df7170456364c8c4d89d16f06da1e787939244f7

Observation 6de38852-8989-4b01-8495-8a41ab86f89d · outbound

This paper cites This setup evaluates the model capacity in concurrent action predictive reasoning for robotic manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in concurrent action predictive reasoning for robotic manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:39.987746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:39.298648Z digest=sha256:683d3cc0986dd0f6b3672e88076fb436f197438708e954d74744f1b5f420824e

Observation dc74911f-a30f-40e4-8d70-b6e3e09c7389 · outbound

This paper cites yes” or “no.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation yes” or “no

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:39.809002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:39.392445Z digest=sha256:f5f1aa31aef8795131b5f9843a64790f86f636b224564249fcdb2f8aa756843d

Observation af3719b3-b7c1-466c-8f02-660250ebc0c0 · outbound

This paper cites John Wiley & Sons Hoboken, NJ.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation John Wiley & Sons Hoboken, NJ

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.324343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:36.849611Z digest=sha256:935bd48dbacc9b9819ea1132ba8c10039fd2e521b34748055c2e57dacf2ed9dc

Observation cab77872-0d73-4b5b-b584-3375f1655071 · outbound

This paper cites S(n) t to denote the whole relationship between n component states and complex states.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation S(n) t to denote the whole relationship between n component states and complex states

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.816986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.758890Z digest=sha256:12b81517b54a66c61ba7816e3627bdb67f0eb8a03ca4719a7b34dd0bb1f6e458

Observation 399c084c-8c05-4a5e-b08e-811189840f70 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation GAIA-1: A Generative World Model for Autonomous Driving

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.281149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.281149Z digest=sha256:2633e0439d64186bdecae8ba83e89ad9d497b490e9aebc51d54041066c21a1c2

Observation ee900967-a049-4644-9de2-b0cda7747bd3 · outbound

This paper cites Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.339688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.339688Z digest=sha256:a37234e257a06cf70f43ea258ccdc7e3450a01054cbe7959ffe944039eec4042

Observation 2f2ff74f-0605-4c25-919e-08b5cca1b525 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.559873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.559873Z digest=sha256:5734d1d74651c23a1b16a35c908b9833bfa3e0affe601359bcb21da14f87ec2a

Observation 7f3f7446-c88c-416c-a422-478b4e23c352 · outbound

This paper cites SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition

Reference 2019

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:22:39.582059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.555217Z digest=sha256:98362b839f1eb8d9f47262cd0f41a0bd25f96311c81992d6f3efa4026904c3c1

Observation 890028db-ee63-46c2-ab8e-277326e274e3 · outbound

This paper cites In Advances in Neural Information Processing Systems, volume 33, pages 10514–10525.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation In Advances in Neural Information Processing Systems, volume 33, pages 10514–10525

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:02.907075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.101582Z digest=sha256:e1315d2ce9f1e9f83b421b66d270f80b681fe2215ece48cf450e689fba2f2253

Observation 6c279471-d033-4ad8-8d46-16a9da93dcc7 · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.996473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.996473Z digest=sha256:826755c22c3cd1df41f223bdb19811eb04809a6f1c33c4a4f65e624f4b0bc5b7

Observation bc149e21-7257-48a7-ad3a-38db00cca78b · outbound

This paper cites Faith and Fate: Limits of Transformers on Compositionality.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Faith and Fate: Limits of Transformers on Compositionality

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.904449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.904449Z digest=sha256:0f43f32882cfe20d8d80965074b950efba6c927e9833b88c4d1405012a77c5d0

Observation 802dd24f-d6f8-43b7-b55e-d9e7c290a502 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation How Far is Video Generation from World Model: A Physical Law Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.307022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.307022Z digest=sha256:d4739773cd3dce02e845e19592ec078b1350fae230309adfe716326e4cff0628

Observation 3392a115-99fe-46c2-9cc2-5a7dc56cf3e8 · outbound

This paper cites In The Thirteenth International Conference on Learning Representations.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation In The Thirteenth International Conference on Learning Representations

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.942959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:22:37.626475Z digest=sha256:3e6913c8735383fc25b620a4c15712a1792b1451cf101b1bbb0f2df9ba3996ad

Pith citing papers

Observation 770e4f03-85ca-428b-9829-82fbbaaf69ce · inbound

Egocentric Bias in Vision-Language Models cites this paper.

Egocentric Bias in Vision-Language Models Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.145985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.145985Z digest=sha256:66907bddf1fae44eaa0f6b9b9e9f54c02abf9f76f2b959bb3fcb67668c57d48a

Observation 17d4b390-cd08-4e15-a2f8-2390ef800198 · inbound

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics cites this paper.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T21:36:42.378149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T21:36:42.378149Z digest=sha256:1685a213070dee003a055a85646df5068e3931536d60b1c24d1333b0b93f7586

Observation f6bf9e4e-5780-48d2-9252-45f48f43f995 · inbound

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning cites this paper.

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T07:27:08.281893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:27:08.281893Z digest=sha256:17146972ea296c9aa3b3bf79327982664362ae1ae8862a71db80a6734072a4cf

Observation d023fc94-0bd3-4f98-b89b-fe4b2cc8c1be · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:36.966727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:36.966727Z digest=sha256:e2fd6c577cbd4959562e93e43ff24ddcfc4d6b416038eb449fdaa9cb9d48586e