Pith. sign in

Paper Citation Record · LEDGER

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

As of 15 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 7 inbound Pith citation observations for arXiv:2508.02095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02095 v2

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:12:38.498012Z

measured 107 of 107 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:33:28.342219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:46:39.977556Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2cd9df5c-473e-46db-9a0e-1eefb1f18b86 · outbound

This paper cites write newline.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.550545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.550545Z digest=sha256:7fe2cf6cdb1de5a2f22ecef94bb83177fe5c505cb6801d4efcaba3b852e22e96

Observation 20cf6149-3466-4b5c-bde3-18d813296273 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.650096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.650096Z digest=sha256:5acaf106ed276c708a34124d91aa64982f51747c299e2c84d861ea81852e0f9c

Observation 03157756-6f0d-434c-bf3d-df149558d8b7 · outbound

This paper cites Phi-4 Technical Report.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.803287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.803287Z digest=sha256:12e0dbfde261a72722d497fcbad73e02826dcaed4fb7dd6c046dbcf50fcb2dd4

Observation da41d6c2-eb7e-4e3f-be99-4e02b8ae4d9a · outbound

This paper cites Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras, 2025.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.926921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.926921Z digest=sha256:7396d11e32c26fff772da196be8b627c7c96f53279efb458bb8898fdb655e740

Observation 023a61aa-dea6-453f-911a-67a3005c28a2 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Cosmos World Foundation Model Platform for Physical AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.027479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.027479Z digest=sha256:8124308a0597385def11b0b0b61b1255680e727b8c187d1de39b72b4600bf160

Observation 664cab2c-53c6-434f-9e6b-512bfa8ac78f · outbound

This paper cites Pixtral 12B.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Pixtral 12B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.127405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.127405Z digest=sha256:6cd674cd240c5df720b99623d716b749828384d41933164f7bd9d03aabb23845

Observation b25dd5a9-5c12-403c-9ed8-387fef226c4f · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models System card: Claude opus 4 & claude sonnet 4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.246355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.246355Z digest=sha256:4b0e2fab54b2ff113b1690f619742285c2b8dca2d6241aa6048393a36a11eae2

Observation ef7effab-fbbc-4fee-8e0a-f5b43ec98a7c · outbound

This paper cites Qwen Technical Report.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.350271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.350271Z digest=sha256:b8eb8344f1058db21a69997070fc9504267e458bdef42e6fb6d6aeb48ad5395a

Observation fca597d7-1351-4284-8a0d-39b77dc38e1b · outbound

This paper cites Video generation models as world simulators.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video generation models as world simulators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.427910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.427910Z digest=sha256:f77ae89cb92d0f8444ee15f28a567431092b6d78c9ffd7d3dde67df23fd7cb5a

Observation 0830ce38-d641-4fc2-970c-412f52cebc59 · outbound

This paper cites Language models are few-shot learners.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.517145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.517145Z digest=sha256:9c124391dac2cf79611d4a4223c61db0b83d90cf5837ff7c1ff558d3a932b86b

Observation 73dec874-a00c-4e02-83d3-17cd5c77c48a · outbound

This paper cites Spatial memory: how egocentric and allocentric combine.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatial memory: how egocentric and allocentric combine

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.611205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.611205Z digest=sha256:4d31f1da51a6ec92dcebb4ebf3383d40a83912b4e2b194fb3247288497acb9d9

Observation acb75277-81a3-4e7c-aff8-73480e991858 · outbound

This paper cites On learning mechanical laws of motion from video using neural networks.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models On learning mechanical laws of motion from video using neural networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.688452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.688452Z digest=sha256:ab0e67342d4a6b335ba7420e61d283b6be1516cb97bf47064151ec6901c63457

Observation 138e8c17-41e3-4474-8df0-a8580409f448 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.764104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.764104Z digest=sha256:279e4c6272e95b3d0ba8d782ed0f6509e2ab97bb6acc531acb4694d68ff4f814

Observation d27a6624-6e91-456d-86ac-9850d9b27fb5 · outbound

This paper cites Visualgpt: Data-efficient adaptation of pretrained language models for image captioning.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visualgpt: Data-efficient adaptation of pretrained language models for image captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.860941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.860941Z digest=sha256:be82932babedb5f07ea12599f8a07d8049fdbbbb3662c982199f3dea830710a9

Observation 6af65b91-3c46-4259-b5ff-e5b2e8fdbee5 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.963176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.963176Z digest=sha256:deebd005505709be5f428dfcb2fe2fe91076c40ae21dd1270c02aa12d1a3d9a8

Observation 7255e154-2e35-4724-8474-3cd9e6b63110 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.079542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.079542Z digest=sha256:ad0c733db6efe754bccd89b4038ad0bea6cac084d74acfbed0b075fdc206406a

Observation e653a79c-4372-4156-9274-87c14bf4e636 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.185606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.185606Z digest=sha256:b1cc4e1fe5bdd92da0ec275452ed41d1a91d2182548aca2c1bc6a90cf1a8799d

Observation 216bba94-d7db-4cb4-8695-c53635ec6aad · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.291359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.291359Z digest=sha256:ec646e75016b88706c9a3321d1476aba5c536c2e33b63c34df36ee748212a121

Observation 9090c138-07fc-4ccb-b204-1fd92cf0aabe · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.405270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.405270Z digest=sha256:e2382c2cb35be2ceb529a5229904bf53ddfbe33dc5a8f120c47c8c667a546ecf

Observation 0f4cdeac-312f-41c5-b891-9a6a14f63b6f · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.493630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.493630Z digest=sha256:6802c6870b89eca2af845c8bb390be6b1d8130587ad5473cf9197dc8ee2462c3

Observation d8a1a71d-81a4-43bc-a62d-bbe78f149e65 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.608958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.608958Z digest=sha256:bd87d8a74ea23afcb74a6efb4fabaf8237cb0d2cce6a4c0297161498931b4920

Observation ea616791-2452-45a3-978a-2594d6df6fee · outbound

This paper cites Myers, and Anna C.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Myers, and Anna C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.740596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.740596Z digest=sha256:545d4c44710a051173ee487beafc754a243d68c44d78cd2d6015e8f690c489cd

Observation e41f1e65-2363-4bad-ac67-c0680f910318 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.832680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.832680Z digest=sha256:d311ede7ee2d95769c5c00675c758e07c86163c83ea7d0295cf48130caa96c70

Observation 0dc3a3a9-09a8-4e53-8863-7f09d3716f67 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.911291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.911291Z digest=sha256:61101122bb4bd33f2c8cb4c1046c9cb505ff432467bd145a952f3b25cfab4c31

Observation 378e8e72-bcfb-4c50-ac30-624e52751d48 · outbound

This paper cites Palm-e: An embodied multimodal language model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Palm-e: An embodied multimodal language model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.027237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.027237Z digest=sha256:c04e63a32294eba3aae8570fdef440e47dce96c7135445685835c479eee3085f

Observation a4be8d57-ab48-46e1-818a-b13203b91d37 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.113371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.113371Z digest=sha256:61d9dafb6ac5c132d1c86e6739e06d2c9f5c1748a50949534d6123f45d518588

Observation 92559efe-a876-4d36-977f-ad9c8d05ab43 · outbound

This paper cites Large spatial model: End-to-end unposed images to semantic 3d.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Large spatial model: End-to-end unposed images to semantic 3d

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.205884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.205884Z digest=sha256:17bdf38645c51fc913e8654b37f29363e8106d214ccb14ce858b328fb4b09f8b

Observation 4b5e0ebd-133a-4469-b99f-ee3463fcf341 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.288256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.288256Z digest=sha256:49a90aae6602f235641c2d7b782c27dfed10ef99bed356727dbda47652b65272

Observation 6f6d57b7-444e-4382-91a7-24bdb6bbe574 · outbound

This paper cites Freyd and Ronald A.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Freyd and Ronald A

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.373301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.373301Z digest=sha256:4944edbebf210cec3bbcf943da6926df4f165e3bab87d30fff58d402ec6c316d

Observation 7b6c2ad2-e003-4f25-b766-b51bcd6c89cb · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.476769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.476769Z digest=sha256:9ff6617c780c8ae7cd1337fc8ef28f413ceda3debf82fa5d3ef839ee50147cd1

Observation c78cae08-3807-4eea-a7d4-d3b6817c0802 · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.543530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.543530Z digest=sha256:3731d90a02aa897b342736ec01435e216ca65917a2c6cd045d371a703d1d5cc7

Observation c873d276-6d9c-4465-aca0-e444a94ae04b · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.629056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.629056Z digest=sha256:5c55fb139cd32da705e720fdd2727d337db5d0dfe26d30c2cb5c58c8bec4f7f3

Observation 76488af5-b9c9-4fc5-b18a-ea10b773592d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Ego4d: Around the world in 3,000 hours of egocentric video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.691518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.691518Z digest=sha256:85de56f31f11151004c27321490846c3af1572bb980cdd3cf6ca6b31c9155225

Observation 70c6eb71-8bda-484c-a9a4-72e10f1ef3d5 · outbound

This paper cites MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.789179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.789179Z digest=sha256:4c2b1b92a13b772b13fda392f7457f06b1c43776a9c37c7995315fd0d6eab675

Observation 768265ba-529a-4ec3-a474-af06077ec43b · outbound

This paper cites Mojito: Motion Trajectory and Intensity Control for Video Generation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mojito: Motion Trajectory and Intensity Control for Video Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.880167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.880167Z digest=sha256:e24fc261460dd3245c2f27238d60a3a0f8fa886372ff61c64eeb69644df7633a

Observation 47e1ad44-514d-4fd2-acfb-2109865dbaaa · outbound

This paper cites GPT-4o System Card.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models GPT-4o System Card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.980532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.980532Z digest=sha256:1a860d33e88ed02446fa8f1cc35bdd60c21a997c1cfadc236cd782731e05eb2c

Observation 91eb9c84-dedb-47fd-a964-c26841f2dfdd · outbound

This paper cites Visual perception of biological motion and a model for its analysis.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visual perception of biological motion and a model for its analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.126549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.126549Z digest=sha256:009de89f3a07ef621fad011329ce07d3ebd2611de8a72318bd397406952e3628

Observation 3926e618-9310-446e-92b2-7a6c75d3ddf5 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.219228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.219228Z digest=sha256:c0682f6caef40a426621048d41375f8af99dbb9e3c67498b7e4e95697a4b4380

Observation c8aa67fb-5870-41f0-be5b-d5ec4fbcb5f7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.312212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.312212Z digest=sha256:346fa2bad8dde69042bb537685efe35fc7fc29558c67a5baded6111f873182a3

Observation f1c06351-8938-4948-8fa7-03c4a5aa02e3 · outbound

This paper cites Decomposing nerf for editing via feature field distillation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Decomposing nerf for editing via feature field distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.363639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.363639Z digest=sha256:14d4f5703e3e1f947c5e508ac0a45d58be483e7fd7a0b5284d8567b5b2f8aefc

Observation 877b0cd2-1373-43a1-a3ba-72745c29571c · outbound

This paper cites On space-time interest points.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models On space-time interest points

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.436299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.436299Z digest=sha256:b14aab175250531db02f29ded95e694acbf4b616b9136467d4d0ece48793d789

Observation 6ef2fe79-e9c0-42e9-b806-8a54f417adbd · outbound

This paper cites Denker, Donnie Henderson, Richard E.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Denker, Donnie Henderson, Richard E

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.519271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.519271Z digest=sha256:f7d696c17c4cf5fd335a684f6f7e7dd22880061d7ba3a4b81384816cab63c876

Observation e19da999-8d1f-4410-81c2-e18f0e18ef1a · outbound

This paper cites an unresolved cited work.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.594550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.594550Z digest=sha256:36538a0fe100ae4ecb9f5faffc3d264a01052efd0e13dfff38fcc5d870a17dbc

Observation 431ecfc8-b97f-4732-9738-36b17afad5ce · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Seed-bench: Benchmarking multimodal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.533498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:35.678091Z digest=sha256:7edaddedb8a4e9a402cdf2b80c08e0eea9abc35716123b0a9e6d616d42506e19

Observation c4556114-2f60-44f3-887f-a6f9b9ad2e81 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.776779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.776779Z digest=sha256:b36f0fc5b08bf26615eebebb11c657d38ef593a08fe79cf5f85b2f333ce28cfd

Observation 967cacbc-4c71-454c-96d2-8db4b59de9d3 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.830990Z digest=sha256:f6034a17ffa89d5967855e5d9bce7f6889b2362d6f86bb1e2533cc93865394c1

Observation 90ee9794-d3cf-40dd-a837-5ce11f4ea952 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoChat: Chat-Centric Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.926537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.926537Z digest=sha256:0ef38842275919bd456d5ff746dacbc5743c91ecc6b4f451eacff08427a39115

Observation f45c45fc-5204-4e98-a655-3cf0b5ec2dc3 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark, 2023 b.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mvbench: A comprehensive multi-modal video understanding benchmark, 2023 b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.521937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.042657Z digest=sha256:21cf56261a8717412502daaf46d857f56ef106564100e08547f361a1fa186602

Observation 0c106f83-17d9-4aa5-a423-d32b6d1a2bdd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.510515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.146633Z digest=sha256:ba6ddc42918ccc6f623414b266353aab6e2ccefffeeda0e513b5b6335d61e046

Observation 9385a981-cc8d-474b-95bd-8d66dcdb2bd6 · outbound

This paper cites 4k4 DG en: Panoramic 4d generation at 4k resolution.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models 4k4 DG en: Panoramic 4d generation at 4k resolution

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.499453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.253992Z digest=sha256:8d7466a73ecf432f21f53c89987826b8c326735592ffd2d33cc5ed2a4715375c

Observation 018d2dbe-f6e4-4b6e-8b39-75981ee43f78 · outbound

This paper cites VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.307349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.307349Z digest=sha256:280ee59de52d8411d47d3e6ce02aff403eb4c6e76e9ee34d1041828d77a44684

Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · outbound

This paper cites Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.368585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.368585Z digest=sha256:980567ffbcf104631c9d9d2abc9a28ac75aa0f37228c16336ca13b9ad2900abe

Observation eed4d6db-0496-461b-b66f-3ea999ddd1be · outbound

This paper cites Visual instruction tuning.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visual instruction tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.432646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.432646Z digest=sha256:567ac0c0b53aeefe07baaf98310a1a7af5ce178c1ec00deb07ce79ffc5a5a644

Observation eea025cc-1422-4aa1-9fc7-c6bc4b033163 · outbound

This paper cites World model on million-length video and language with ringattention.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models World model on million-length video and language with ringattention

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.481792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.521405Z digest=sha256:0b804a0b4f0a032b3e11984112d81581ad1989ca627c16504681557fb1bd25e1

Observation 51dc2eca-82f2-4fda-9721-8dc7dd0f5e69 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.616482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.616482Z digest=sha256:a64113f1c652b631afa8a54d8fcbf8603f8ca58fdc7624f5cca5e22a650886c8

Observation 916677d5-7baf-4fd4-8920-2e4d2d5c9dfd · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.470085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.726798Z digest=sha256:fb230b44e8bb9e55d74d5a4837f091186334264ec562fa78e9f034713399516c

Observation c2891cc0-810f-47de-a0ca-3fdf9bde578e · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.834649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.834649Z digest=sha256:cc4911e7a3eba7dda97d5ff6dbef03af4e03cce2956ab696640421700671aba5

Observation ec286d9c-b5c3-4e5f-8d79-9cf591fb5f30 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.458973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.958887Z digest=sha256:1cd173094ce29f25c12ef1ca9b29abaa2185361301fe06910556dfb6ceaa6d09

Observation 0e62240e-473e-426c-a6cb-fcad01a37c68 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.037955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.037955Z digest=sha256:df389fb16542d1c91ab6b2f1443ca136ae21fc0094276153f68ed1db2642e0eb

Observation be8ee27b-effc-4a6c-a1a1-ab9a57d7f7fd · outbound

This paper cites Marr and S.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Marr and S

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.446426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.137351Z digest=sha256:5d00bd90e211a3bcf5b6a2b9ae41fa6eff617f971ec05538ebe07bef081f9368

Observation 6ea2a558-96e3-45a5-b926-575bde0fd5f7 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.434678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.223152Z digest=sha256:55e57458cb35104a1f420c8df9f5488a784ac8038c0b42078ccbcf2f89ec4ad8

Observation 67684500-24fc-4457-a575-c13145b72f96 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.260203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.260203Z digest=sha256:e3bd81413ce3ff891ddaa2ed092cf256fb9486e7f72fb769a9f1dceabefc1fa2

Observation 32b720db-3f87-47e2-b9ef-84a16d44b79a · outbound

This paper cites Sceneteller: Language-to-3d scene generation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Sceneteller: Language-to-3d scene generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.421615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.263896Z digest=sha256:6155512ac87c3cf1ddafeb810768ebec0fb44c61d42df94764031b4697b1c59f

Observation 00ff1076-2416-46ed-9b02-321ced9abc42 · outbound

This paper cites Hello gpt-4o.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Hello gpt-4o

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.408122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.332427Z digest=sha256:45579bf9cae79104a0a778909310a9286320315b63e492c6e3a208b97f97db39

Observation 16a0106c-88aa-4b60-a40b-59df9e41021c · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.457514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.457514Z digest=sha256:6e168cb1ab08f867b23261e786b90d71665e04f15a69f4e482306d3e9961deec

Observation cb64b7cd-bf99-476e-b198-b267123171ca · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models A benchmark dataset and evaluation methodology for video object segmentation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.395023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.521570Z digest=sha256:76c1e352b4fff0a3d9cba74155500c36047f32f37048a8e81bef6dc829fb5b9b

Observation 8fc0dd45-2e18-43e2-b8d8-060dc2ead32d · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models The 2017 DAVIS Challenge on Video Object Segmentation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.630315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.630315Z digest=sha256:ad330a428a9e03cc7cd591c75a46cab39cf13635dee675cd07f7cce66e57be77

Observation 1491ac1c-17b6-4d41-80bb-7c8c12409df4 · outbound

This paper cites Improving language understanding by generative pre-training.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Improving language understanding by generative pre-training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.734832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.734832Z digest=sha256:3f933091c49b28d3249edf8492a2b825143b59449bc6d2a2bab5961ac22a67e1

Observation c5f338b6-8a45-45b4-b407-6057abf106f6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Learning transferable visual models from natural language supervision

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.861862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.861862Z digest=sha256:4856e63ceb894f898776621b42f36708805cac9a8d4539dd28da5d5b4fe9ae52

Observation 09c004b0-1cea-4029-9754-68e1f4137c20 · outbound

This paper cites Learning to localize objects improves spatial reasoning in visual-llms.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Learning to localize objects improves spatial reasoning in visual-llms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.368011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.012050Z digest=sha256:6b2fe4732fcb1879df941e4a139898dfbe661c6f1669d8b7b97d92e1e5f562bb

Observation 4b53dbe3-ed92-42eb-ae97-181e0f89d6b8 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Two-stream convolutional networks for action recognition in videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.355662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.140748Z digest=sha256:cd4a19acf6cbf7c5d85ae532b1e146b9afc0de3766ec8e5fc2083385cbf15522

Observation a5d69fda-8888-4ce4-8f7e-f77463128fea · outbound

This paper cites Spelke and Katherine D.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spelke and Katherine D

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.343276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.248287Z digest=sha256:64fc312e2ed838a058af4783dea47f2c55e7dc4d03c8b674584565b8d5938109

Observation ea651a9d-1b27-4bff-8db3-46c22c15aa32 · outbound

This paper cites AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.341363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.341363Z digest=sha256:7ade549b88de0dab36e78fa4447898a1a3ca50ea458b7561b2789395aab9a5b4

Observation 2a2386b2-d0ec-4ed5-ba0b-9cc7e2f16731 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.401698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.401698Z digest=sha256:1fcb050a6b117022b545ab090dc2e7c95866b38da5fa15ef5a71ddd7df283c06

Observation fb9915e8-04f5-4ca1-8b26-eb2b4cee5e84 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.411315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.411315Z digest=sha256:9096b79fbe572a697cd75c02b28f080a2c57abc85b875dda4771c4585a291367

Observation 9945a242-91fa-48b0-b0ef-6ae9974d16ea · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.419175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.419175Z digest=sha256:f6b451f06f6735118d49f1163839ded39c2fdb2b2db14a2ad063179b55b08217

Observation 464dbc8e-7267-4361-8499-dae2e896ccb3 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.422740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.422740Z digest=sha256:89e31e37726d813af5d628c7e57ed375ec4e3c3128ee988e3f871e047269f6ba

Observation f8c0ae7d-3c59-4167-ae14-2db5dc8abad1 · outbound

This paper cites Action recognition by dense trajectories.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Action recognition by dense trajectories

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.331233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.426524Z digest=sha256:aeeb87b2d3b2e888fbfbedaf173f659198cf5cec8e115113b50f5fcbe7c9796d

Observation 13f125b0-b270-4d4e-b194-a727aecc27fb · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.429768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.429768Z digest=sha256:dcc58ed4c48a9e51d89253e0461f5cca30e4244f83cddb6165ed0ab365199e51

Observation 018498f3-03d9-4b86-b582-08f36c393dc4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.433620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.433620Z digest=sha256:90b95c58647553a9348f41def3eb698dd84c7db0e64beeb8dec6973e28a7a247

Observation 95d03c2e-bec6-456a-905e-5231ad6d9159 · outbound

This paper cites Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.437496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.437496Z digest=sha256:34949b951d4188c8d1f2486e7391ed44ea76aa0940ed4280cbb8b1526782665d

Observation 3f667929-c395-473a-ade8-bd8a3d499292 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Internvideo2: Scaling foundation models for multimodal video understanding

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.318380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.441640Z digest=sha256:01a7e2c0c665f5ff7981a98c72b1c454fcb79da82ca864ec64d73100e9ce51f0

Observation 0e018751-258d-4358-8242-05dba5a50bca · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.444811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.444811Z digest=sha256:aa4cb4401ede8c381d050afd5ae58f67cc51a7d6b1c87d4372cfd253841ff32d

Observation 5650af5a-e4d5-43aa-8037-9f400c9a7c0e · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Finetuned Language Models Are Zero-Shot Learners

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.448025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.448025Z digest=sha256:35b9d3c799130a6b756edec79601d1c7bdfb693e53961b0b3533c8ad6411d192

Observation b63d73dd-d739-4c97-b02d-66b3c0a5b79b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.451316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.451316Z digest=sha256:53e8409e772664c7ca6c400d64ed0fa12883cfb3a3a359d06befb01de450908b

Observation 45e3e534-fd64-4804-86fc-5ea62873d6f6 · outbound

This paper cites Cat4d: Create anything in 4d with multi-view video diffusion models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Cat4d: Create anything in 4d with multi-view video diffusion models

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.298636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.454660Z digest=sha256:32de25df1d2e9d5d3127b6c0c0de2436fbe902037e27e92d19c1796de02a6969

Observation 2e0bbea3-32fc-46ed-97e6-3782efc20926 · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.286966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.457594Z digest=sha256:65d4cf261620158f8fc8da52a889148cea493058e302f1c011d4788a94cf2cf7

Observation 9c235fba-8e0e-4c66-b221-1238e47e675a · outbound

This paper cites Grok-2 beta release.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grok-2 beta release

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.276520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.460780Z digest=sha256:c0217680643f06e97d394a2069bcfedfddaf07f691be36b2a51a7ca6d3e3b0d0

Observation 78578b2e-8818-4370-b962-4ba188d86139 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Youtube-vos: Sequence-to-sequence video object segmentation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.265029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.463793Z digest=sha256:dab54e28277cc73ee2187b67a16c19a41f46bf9b0db009bf6f1fea4aff50c422

Observation bed3bb16-6535-4718-8263-7575eca19771 · outbound

This paper cites Qwen2.5 Technical Report.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen2.5 Technical Report

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.466933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.466933Z digest=sha256:3c2f47f00b7cf19ca39b569c8b14143f49358f77262ba3a77efbe64b208c07bb

Observation a2b826cd-e5d1-4f83-9b2d-45848f42c01e · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.469833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.469833Z digest=sha256:168e8e55b13c580cb6f5684a36f898bd20e59a04415a01f54cb72f4cfb9c7f46

Observation e178e433-1612-4a94-ab5f-f9e3953d7358 · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.472814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.472814Z digest=sha256:3a887c8df90e0baf2549ff62a5d4e2569c2ea24e6bf25127800d826b69e124ca

Observation 3778f7a8-6891-4d1a-84fd-f2edad6885a4 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.254346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.476387Z digest=sha256:9ed44339693839a048edad76f75852bad63737a941673a13aeef137b03a7e85a

Observation e68b6a35-fab0-45ca-a137-d16da4d723e5 · outbound

This paper cites Improving 2d feature representations by 3d-aware fine-tuning.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Improving 2d feature representations by 3d-aware fine-tuning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.243822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.479189Z digest=sha256:b1cf8c96f26a1c2163c8212c00c1856c6227a1ed9baa38b0de897fb4848edfd8

Observation 906c1777-421d-4184-888d-adc06cbdecfd · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.481827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.481827Z digest=sha256:f4b65a92c8c827521e36310878963069c2f0428a688d928af8825ee5ec2f4d72

Observation e8ab4673-66da-40db-9843-af89c382160e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.485190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.485190Z digest=sha256:e30c0fc8831f11f56a74f68aa2c73e01f3dca6db9024df1d054b70cd20d48fbd

Observation 7c9db707-929a-4ec3-9346-a4319d2da646 · outbound

This paper cites COMBO: Compositional World Models for Embodied Multi-Agent Cooperation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models COMBO: Compositional World Models for Embodied Multi-Agent Cooperation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.488149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.488149Z digest=sha256:32a3bc946a2334133a9b251326be3103055a00da64585898be5ee95d62d72646

Observation ed33a6e8-ec2f-497f-b9c9-448ad37156f5 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, 2024 b.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Llava-next: A strong zero-shot video understanding model, 2024 b

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.231812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.491581Z digest=sha256:84e4ac2ff1c8a508519e66250875aad6d9b45da2d3aa0c98e7748ba41e8b50d4

Observation bf9ba7c1-8894-43ca-9c45-32e62ae6b963 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.494969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.494969Z digest=sha256:b38e1a0ebc35e98a29f242a6c5897dd6fcc816b3defb59db76f0fdd76b6bf7d2

Observation eb72600d-c8fd-4187-b371-bee4d690d62b · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.220174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.498012Z digest=sha256:3b7075057de2afdced90f6cac219b6825541cc49a2904e9ae226d78759d24e96

Pith citing papers

Observation 52903efc-fbc8-49af-b197-4bdc8b134722 · inbound

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km cites this paper.

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:37:41.277166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:37:41.277166Z digest=sha256:a9c0b9601f752f9f72f9c41c83d79261b6301405a26ec34b791cfc9d222933ee

Observation b9d7bddd-8785-4d5c-bb5e-1031eda8cec5 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.883037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a86a76f4dea3e39320116eb712cfa8923128b0900020380c59b3a4179cd4a428

Observation b428e44c-bccd-4e54-903b-c1e3f9ef82b3 · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.226528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:c1258115a359dce37118a038244769a20da77b8a96ff9b03cd347e05b1a1b472

Observation 261994e6-167f-4bdc-b274-17dfea95c1c4 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:39.980679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:0aab726f70cb727c1cd23c09801b91d4da490cda3c3d433699436c17f22db1d1

Observation febf3a8a-1ac7-4b17-b709-d0e8ce2faf8b · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:ff11ea430ea1badea43da93104e784c3148b41b6cd69f6caab85246b070c6137

Observation 493af086-78f3-47ac-ba1f-8cd2174e622e · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:52.231964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:52.231964Z digest=sha256:fe51187c4f632631c738aa9a7f173d32048bb4ddfdfee92b41b9610549a4ddae

Observation d9ce87d5-a09d-4627-9a8c-8812128504eb · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.342219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.342219Z digest=sha256:23c82ce57289a9a7d19928f658649d2854634a253ef21d45bee6005980dd2518