Pith. sign in

Paper Citation Record · LEDGER

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 6 inbound Pith citation observations for arXiv:2508.02095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02095 v2

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:12:38.498012Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:37:41.277166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:46:39.977556Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2cd9df5c-473e-46db-9a0e-1eefb1f18b86 · outbound

This paper cites write newline.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.550545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.550545Z digest=sha256:3df061a6a62322024a9c085c99afe90c4c2353ce447203ef162bc41625961341

Observation 20cf6149-3466-4b5c-bde3-18d813296273 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.650096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.650096Z digest=sha256:7724f26f243ae7a92f7b215c8ffc8ea315fce33f178fa109e54ad731f252156c

Observation 03157756-6f0d-434c-bf3d-df149558d8b7 · outbound

This paper cites Phi-4 Technical Report.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.803287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.803287Z digest=sha256:78a31e78c1f2482978d4aa512b47af994ba331348b66a3071d7cd9c0893bbf7c

Observation da41d6c2-eb7e-4e3f-be99-4e02b8ae4d9a · outbound

This paper cites Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras, 2025.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:31.926921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:31.926921Z digest=sha256:562054bc186c0063963c7a875ca252302b457cfacc42b08f26abd091319fc5eb

Observation 023a61aa-dea6-453f-911a-67a3005c28a2 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Cosmos World Foundation Model Platform for Physical AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.027479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.027479Z digest=sha256:a261017239ef6cd3f1527d017e75c1d891db26843ecf929395c490d36485d1ff

Observation 664cab2c-53c6-434f-9e6b-512bfa8ac78f · outbound

This paper cites Pixtral 12B.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Pixtral 12B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.127405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.127405Z digest=sha256:2d0fb28e2f2cabb71781d2f5bb25c47ebcb000af253600b46bfd034a17696370

Observation b25dd5a9-5c12-403c-9ed8-387fef226c4f · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models System card: Claude opus 4 & claude sonnet 4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.246355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.246355Z digest=sha256:f84b443a924dbd6171a0b3922195f2c5172819e3fe39058950c85f954a2b52c6

Observation ef7effab-fbbc-4fee-8e0a-f5b43ec98a7c · outbound

This paper cites Qwen Technical Report.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.350271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.350271Z digest=sha256:82cbf1023fb5120099b9a6b788f083ce7b43af14cef716543b2356546bc29e8b

Observation fca597d7-1351-4284-8a0d-39b77dc38e1b · outbound

This paper cites Video generation models as world simulators.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video generation models as world simulators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.427910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.427910Z digest=sha256:33640948915381878c224e65df3441848680b573abe810423d1db4fb474b8ace

Observation 0830ce38-d641-4fc2-970c-412f52cebc59 · outbound

This paper cites Language models are few-shot learners.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.517145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.517145Z digest=sha256:47f4dce26a42c6145dbe1f61b9ed1f612bea72e20498703811e5cff724015dc2

Observation 73dec874-a00c-4e02-83d3-17cd5c77c48a · outbound

This paper cites Spatial memory: how egocentric and allocentric combine.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatial memory: how egocentric and allocentric combine

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.611205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.611205Z digest=sha256:f199c47cac3c027fb23170fa2d1b2c71c6d766e0748557fb36e54edf8dfd4729

Observation acb75277-81a3-4e7c-aff8-73480e991858 · outbound

This paper cites On learning mechanical laws of motion from video using neural networks.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models On learning mechanical laws of motion from video using neural networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.688452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.688452Z digest=sha256:1e65191c7e3fd2f8b4e58352e020ad882703d0ea55eaf075ddb176165c3eedc1

Observation 138e8c17-41e3-4474-8df0-a8580409f448 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.764104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.764104Z digest=sha256:1a18ae603d9ef4b6138dcfc06c6c856aa5c02a1c4eda8488d18921f931db34ee

Observation d27a6624-6e91-456d-86ac-9850d9b27fb5 · outbound

This paper cites Visualgpt: Data-efficient adaptation of pretrained language models for image captioning.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visualgpt: Data-efficient adaptation of pretrained language models for image captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.860941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.860941Z digest=sha256:f900f3b6702f3929ebdad7e24778404436e92c101086636af88edd9e757304f2

Observation 6af65b91-3c46-4259-b5ff-e5b2e8fdbee5 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.963176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:32.963176Z digest=sha256:27f2bb9e499266347e0c2d1400e7eb1910751bdc8b38cef4de89061956347734

Observation 7255e154-2e35-4724-8474-3cd9e6b63110 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.079542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.079542Z digest=sha256:20127ff8b63f2d7a7c043842db08b8d37fb428856b6ab83f92cfaff34f32b7f3

Observation e653a79c-4372-4156-9274-87c14bf4e636 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.185606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.185606Z digest=sha256:bb94e30e166c0713288529284c8944922cbe1953a9099c2733f88dcc62b8e4d5

Observation 216bba94-d7db-4cb4-8695-c53635ec6aad · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.291359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.291359Z digest=sha256:78585e32d4e95116e1aac4d2b56485a2942dad3cf5ec884f97cb6c44192ce049

Observation 9090c138-07fc-4ccb-b204-1fd92cf0aabe · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.405270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.405270Z digest=sha256:aafcec0a2f1ede9a52aab0e61e6228eea7155a9b8c8f6f35415c0dbd128d3b83

Observation 0f4cdeac-312f-41c5-b891-9a6a14f63b6f · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.493630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.493630Z digest=sha256:ded7bede6b4be5f40ddb7366977690912b839a5425faff6332455af96d8ad5d0

Observation d8a1a71d-81a4-43bc-a62d-bbe78f149e65 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.608958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.608958Z digest=sha256:e0d2272e77674333b9561996c97edffc8f58812780f7fbf0c75859dee8bbc8a6

Observation ea616791-2452-45a3-978a-2594d6df6fee · outbound

This paper cites Myers, and Anna C.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Myers, and Anna C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.740596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.740596Z digest=sha256:d5808e334b687a5d2beb07900f7bae2b81e7a0d9a992fa9cc85886998f70100b

Observation e41f1e65-2363-4bad-ac67-c0680f910318 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.832680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.832680Z digest=sha256:4231b371726ab6c25210cbc1524f8430e2d79ec3095ca87f4004d2f3e1eae69b

Observation 0dc3a3a9-09a8-4e53-8863-7f09d3716f67 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:33.911291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:33.911291Z digest=sha256:95b0e6230ce7a650109bd31528337f3f7497c7199663e211d80617b308d92913

Observation 378e8e72-bcfb-4c50-ac30-624e52751d48 · outbound

This paper cites Palm-e: An embodied multimodal language model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Palm-e: An embodied multimodal language model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.027237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.027237Z digest=sha256:99dd471cca903c7a40d8f9868c20f01a3f9c5baba8c6405f0d84eafc55f05220

Observation a4be8d57-ab48-46e1-818a-b13203b91d37 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.113371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.113371Z digest=sha256:abe3bf28d90e7987c69272961e41ebeab93f1f30ffa8533f8d7171f7b555c49f

Observation 92559efe-a876-4d36-977f-ad9c8d05ab43 · outbound

This paper cites Large spatial model: End-to-end unposed images to semantic 3d.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Large spatial model: End-to-end unposed images to semantic 3d

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.205884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.205884Z digest=sha256:fb65cf4169f84293e1bc2bcc9f88d44aecc92a2d87688d9b23112a0732355135

Observation 4b5e0ebd-133a-4469-b99f-ee3463fcf341 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.288256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.288256Z digest=sha256:98e46360ad1116f2778569ae61b8cc90df0545ce814fa49995dcc616e22cc542

Observation 6f6d57b7-444e-4382-91a7-24bdb6bbe574 · outbound

This paper cites Freyd and Ronald A.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Freyd and Ronald A

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.373301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.373301Z digest=sha256:c1d6e03be3801f75353cdc36659430dbb63aff5e2c75dd4f0c3f93ea75acd4b4

Observation 7b6c2ad2-e003-4f25-b766-b51bcd6c89cb · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.476769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.476769Z digest=sha256:5646f0fabc50a0f5c94270562817be019ac3bdb374ce2f44ab8cc58328f7f74a

Observation c78cae08-3807-4eea-a7d4-d3b6817c0802 · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.543530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.543530Z digest=sha256:909e5bf1ecdfcde94698db3b56913f1de2dd96a0b50d800983e07c513d357735

Observation c873d276-6d9c-4465-aca0-e444a94ae04b · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.629056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.629056Z digest=sha256:12df9c2a509a577e32b6b4bae268828f8617c9f85001b08a894a6eedc5c2cbdb

Observation 76488af5-b9c9-4fc5-b18a-ea10b773592d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Ego4d: Around the world in 3,000 hours of egocentric video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.691518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.691518Z digest=sha256:007f02b1df1051e7f779b899c347a3bc3562433b0e30595de8f5ad1d0d8e7982

Observation 70c6eb71-8bda-484c-a9a4-72e10f1ef3d5 · outbound

This paper cites MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.789179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.789179Z digest=sha256:005edf50d36244e03acbcbc1c01fdf95320352476c4176fd0645286eb0271d4a

Observation 768265ba-529a-4ec3-a474-af06077ec43b · outbound

This paper cites Mojito: Motion Trajectory and Intensity Control for Video Generation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mojito: Motion Trajectory and Intensity Control for Video Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.880167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.880167Z digest=sha256:66747480950570761c060587b949148c2d13305fc6b893b803e2fc662e75ae41

Observation 47e1ad44-514d-4fd2-acfb-2109865dbaaa · outbound

This paper cites GPT-4o System Card.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models GPT-4o System Card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.980532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.980532Z digest=sha256:e374aa4f5d4c639098bdc4858efc7f34ed3e9f5df628fd6310d585b6e57fcfc2

Observation 91eb9c84-dedb-47fd-a964-c26841f2dfdd · outbound

This paper cites Visual perception of biological motion and a model for its analysis.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visual perception of biological motion and a model for its analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.126549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.126549Z digest=sha256:86e610b0eb48c5847f56b1f580203d196d84a5b5e320b839a20892f15e582556

Observation 3926e618-9310-446e-92b2-7a6c75d3ddf5 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.219228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.219228Z digest=sha256:137c3accffa60f9727cc288d7cece4ea85331950952b3d35dd95b0526274fa8e

Observation c8aa67fb-5870-41f0-be5b-d5ec4fbcb5f7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.312212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.312212Z digest=sha256:902fc5634ac8fd759d3e90ee5b503cff5db8555788fef382eb06a5a7cd47bd62

Observation f1c06351-8938-4948-8fa7-03c4a5aa02e3 · outbound

This paper cites Decomposing nerf for editing via feature field distillation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Decomposing nerf for editing via feature field distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.363639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.363639Z digest=sha256:691b186539cfb6f4b6c7cc67a3eebea18c98feb5b123c02ff68bb0f77de7cd06

Observation 877b0cd2-1373-43a1-a3ba-72745c29571c · outbound

This paper cites On space-time interest points.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models On space-time interest points

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.436299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.436299Z digest=sha256:643975357b90db1fecfc7040982876316313b4ffea256362cab26bbd7215f0f0

Observation 6ef2fe79-e9c0-42e9-b806-8a54f417adbd · outbound

This paper cites Denker, Donnie Henderson, Richard E.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Denker, Donnie Henderson, Richard E

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.519271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.519271Z digest=sha256:95f84f7c77c8e149fe6c39bc2db1976d48a97a239b9722271a0456d7daf5cbca

Observation e19da999-8d1f-4410-81c2-e18f0e18ef1a · outbound

This paper cites an unresolved cited work.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.594550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.594550Z digest=sha256:42ccff6a93dd27bbecd1f8f8671e8f64e2c19bdefb87de2576cfd4c874d24155

Observation 431ecfc8-b97f-4732-9738-36b17afad5ce · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Seed-bench: Benchmarking multimodal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.533498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:35.678091Z digest=sha256:762b74eb99d2be3b31a688e0865579fbf66e3f2298f9a62a5a9f94648455d116

Observation c4556114-2f60-44f3-887f-a6f9b9ad2e81 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.776779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.776779Z digest=sha256:83fb72bd7901939a62bfa636f1f674f5d2c723d01780eefe25ce77abb5b3b755

Observation 967cacbc-4c71-454c-96d2-8db4b59de9d3 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.830990Z digest=sha256:8bf7b480324958e3494d2a7e169465c61ec97516d4642688341e5bfd74ac3b58

Observation 90ee9794-d3cf-40dd-a837-5ce11f4ea952 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoChat: Chat-Centric Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.926537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.926537Z digest=sha256:d00010d72885633839fcc4ef4fbe49b2d8b43ddf14989845b0a808c41463b330

Observation f45c45fc-5204-4e98-a655-3cf0b5ec2dc3 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark, 2023 b.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mvbench: A comprehensive multi-modal video understanding benchmark, 2023 b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.521937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.042657Z digest=sha256:3ecdf6657bc921ff1c8c2c3337daab5c6130d6bb10a12873f39ab7ecfbbaa753

Observation 0c106f83-17d9-4aa5-a423-d32b6d1a2bdd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.510515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.146633Z digest=sha256:7782e29662c38aedba27a2ec09ae5607bcb2de90a6727f91d59782cb7ca20667

Observation 9385a981-cc8d-474b-95bd-8d66dcdb2bd6 · outbound

This paper cites 4k4 DG en: Panoramic 4d generation at 4k resolution.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models 4k4 DG en: Panoramic 4d generation at 4k resolution

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.499453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.253992Z digest=sha256:c19403d3ae2de21fb213c49a5b7f05c181ecf0c0d09c9e3d7773ef2991b85002

Observation 018d2dbe-f6e4-4b6e-8b39-75981ee43f78 · outbound

This paper cites VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.307349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.307349Z digest=sha256:a97ee2d4ac65c5b8773b29d4708645b0a5ed94e460e01ceab44a4a85bd57770d

Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · outbound

This paper cites Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.368585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.368585Z digest=sha256:2277b6b74f5b52d9c59ab41efa140e198fe4fb182acc5aac9e278ca6f7dacfb5

Observation eed4d6db-0496-461b-b66f-3ea999ddd1be · outbound

This paper cites Visual instruction tuning.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visual instruction tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.432646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.432646Z digest=sha256:a4110544e0439deb73bd9cab8157e48760d1e3b31b556a0185cba075d2204083

Observation eea025cc-1422-4aa1-9fc7-c6bc4b033163 · outbound

This paper cites World model on million-length video and language with ringattention.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models World model on million-length video and language with ringattention

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.481792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.521405Z digest=sha256:23dbd6bf1d150c717791c760b8ac712bd65784ce44b207b7a55eafa352124aed

Observation 51dc2eca-82f2-4fda-9721-8dc7dd0f5e69 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.616482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.616482Z digest=sha256:d66bde7b702fab5c68a9b79b0a4c8f1b46011334dd32e1939c2346a900d6c3c7

Observation 916677d5-7baf-4fd4-8920-2e4d2d5c9dfd · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.470085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.726798Z digest=sha256:f8f6b3f3d326342b8bcdb71ae6f25b9616488c09447b7f2f07f31179191bb533

Observation c2891cc0-810f-47de-a0ca-3fdf9bde578e · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.834649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.834649Z digest=sha256:4a29aa8f50f4c38a05fee2de0084850bf9c04bfd05ae27372f5e2d50cbd20d29

Observation ec286d9c-b5c3-4e5f-8d79-9cf591fb5f30 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.458973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:36.958887Z digest=sha256:f474cc8a2574cb5923911ecac0d5d7ee40e8b551ac9f30177909c63d58ef89f8

Observation 0e62240e-473e-426c-a6cb-fcad01a37c68 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.037955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.037955Z digest=sha256:1d891e6b4c09a40bbc0f4b4cfa9dd00469780403dec45f9b7b384fa0bbbd7512

Observation be8ee27b-effc-4a6c-a1a1-ab9a57d7f7fd · outbound

This paper cites Marr and S.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Marr and S

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.446426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.137351Z digest=sha256:20fdc470b58f38ccfdad408b026209e3d67721693b3be6d0040f58643446e2c1

Observation 6ea2a558-96e3-45a5-b926-575bde0fd5f7 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.434678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.223152Z digest=sha256:5524f4f729763f91bfca79017d681299fa861e28943c163a2246ec1dfaa0fddb

Observation 67684500-24fc-4457-a575-c13145b72f96 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.260203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.260203Z digest=sha256:af5316d793032cd50f18eb8aba101824e836ea5f12dba9e2b3d21bc18892851d

Observation 32b720db-3f87-47e2-b9ef-84a16d44b79a · outbound

This paper cites Sceneteller: Language-to-3d scene generation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Sceneteller: Language-to-3d scene generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.421615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.263896Z digest=sha256:7341d2f078dde063a2a16fed6d28f3a546fe18a7d162a7eae43ab1dc3a074c88

Observation 00ff1076-2416-46ed-9b02-321ced9abc42 · outbound

This paper cites Hello gpt-4o.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Hello gpt-4o

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.408122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.332427Z digest=sha256:85c93a610ca55e5d9c8ce825e2c6a270cfd69460bd7cffb3dd39bf8614bd1f82

Observation 16a0106c-88aa-4b60-a40b-59df9e41021c · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.457514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.457514Z digest=sha256:1c8a1729fa44d7f5b1028bd552e0ba07ffd4b3a62603922355a362f875576794

Observation cb64b7cd-bf99-476e-b198-b267123171ca · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models A benchmark dataset and evaluation methodology for video object segmentation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.395023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:37.521570Z digest=sha256:de3052bdbe8e613af7e2c390cf970c6d0626762e5156b1075ed2bac64d170e7a

Observation 8fc0dd45-2e18-43e2-b8d8-060dc2ead32d · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models The 2017 DAVIS Challenge on Video Object Segmentation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.630315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.630315Z digest=sha256:79117e4c3bbc8c573112dcf7eecbafa57613e1e886b4147f4a08de23cd7f0230

Observation 1491ac1c-17b6-4d41-80bb-7c8c12409df4 · outbound

This paper cites Improving language understanding by generative pre-training.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Improving language understanding by generative pre-training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.734832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.734832Z digest=sha256:192550625773aa015f35e3d07e46cef4d2dd1182a53c7c2d24ea81f24bb48020

Observation c5f338b6-8a45-45b4-b407-6057abf106f6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Learning transferable visual models from natural language supervision

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.861862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.861862Z digest=sha256:9b2482415b4b9cb62bbaee98d3d4fc64cce5c543bd1551c518999ea56810a7fc

Observation 09c004b0-1cea-4029-9754-68e1f4137c20 · outbound

This paper cites Learning to localize objects improves spatial reasoning in visual-llms.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Learning to localize objects improves spatial reasoning in visual-llms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.368011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.012050Z digest=sha256:69794e938bfec3652e0a6b2a19d94f5dca09c1331241bfffa3db1d11e8a1f8d8

Observation 4b53dbe3-ed92-42eb-ae97-181e0f89d6b8 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Two-stream convolutional networks for action recognition in videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.355662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.140748Z digest=sha256:7c3c928abae416a124370a9cef764587db273678152b35b1d8733e765a5b182b

Observation a5d69fda-8888-4ce4-8f7e-f77463128fea · outbound

This paper cites Spelke and Katherine D.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spelke and Katherine D

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.343276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.248287Z digest=sha256:03b7987be612d6aeb5db92c47b5272450c9fe2e6f71a733b10b5646fd344766a

Observation ea651a9d-1b27-4bff-8db3-46c22c15aa32 · outbound

This paper cites AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.341363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.341363Z digest=sha256:2c6362692f73417fa730a08af673f206706094ed39beef736ea80b8eec307f3e

Observation 2a2386b2-d0ec-4ed5-ba0b-9cc7e2f16731 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.401698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.401698Z digest=sha256:94a065d896104b679ef537b35386c125736b3784f86b1fb9cef74b7811e99532

Observation fb9915e8-04f5-4ca1-8b26-eb2b4cee5e84 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.411315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.411315Z digest=sha256:c2ba20925f04de517e5aee12df93cf9b4d7390dd17a1087d6817fae91b73c22a

Observation 9945a242-91fa-48b0-b0ef-6ae9974d16ea · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.419175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.419175Z digest=sha256:e001d97fc0158d1c46901672f6b1d4143d87f42184f8d7c60bbfd527cc3c2143

Observation 464dbc8e-7267-4361-8499-dae2e896ccb3 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.422740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.422740Z digest=sha256:ca4286a5d808fe77cd37a59ccaa8a60361a3a92281aaeae09364d0fe229b218d

Observation f8c0ae7d-3c59-4167-ae14-2db5dc8abad1 · outbound

This paper cites Action recognition by dense trajectories.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Action recognition by dense trajectories

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.331233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.426524Z digest=sha256:3664ac9759b778f32a0bc662f1656982fb6682269c71a70dd2d1e56a38a0db2b

Observation 13f125b0-b270-4d4e-b194-a727aecc27fb · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.429768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.429768Z digest=sha256:8ceee03875926d0ca25679571e94186416e5f1e438a80b201fdae38b27ab7fd6

Observation 018498f3-03d9-4b86-b582-08f36c393dc4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.433620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.433620Z digest=sha256:4628452911656afe8a7eef7435c9f03c133230ae9372bc691855b0153cd3d806

Observation 95d03c2e-bec6-456a-905e-5231ad6d9159 · outbound

This paper cites Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.437496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.437496Z digest=sha256:da63e7f08d2ccba29610a8db50330a84723fffd0ded7bdde13be3fcf9aa4f4dd

Observation 3f667929-c395-473a-ade8-bd8a3d499292 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Internvideo2: Scaling foundation models for multimodal video understanding

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.318380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.441640Z digest=sha256:d2b584e607fa6dbd8fda8538c67a4673c008a6e8d1be92581c38eeff13f3a3c3

Observation 0e018751-258d-4358-8242-05dba5a50bca · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.444811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.444811Z digest=sha256:f77a4eb191f1164545ea5ebbc724fb75a222071ad33ad7c4b6e9950f19522756

Observation 5650af5a-e4d5-43aa-8037-9f400c9a7c0e · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Finetuned Language Models Are Zero-Shot Learners

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.448025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.448025Z digest=sha256:3d3d9e46b06d4d0891872e5304c1b27788680f2f32242230749c5021b27499d7

Observation b63d73dd-d739-4c97-b02d-66b3c0a5b79b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.451316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.451316Z digest=sha256:4589f5f8d7181a2150adb1beef4974d8eb1b7f8e06aed2278b0e981e42659afc

Observation 45e3e534-fd64-4804-86fc-5ea62873d6f6 · outbound

This paper cites Cat4d: Create anything in 4d with multi-view video diffusion models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Cat4d: Create anything in 4d with multi-view video diffusion models

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.298636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.454660Z digest=sha256:598404f054c0c59159abcc8426ea428072b473ca0aeb524dfa305ba29b75a16a

Observation 2e0bbea3-32fc-46ed-97e6-3782efc20926 · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.286966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.457594Z digest=sha256:484bc1630ea5a6d0c6406bfef1231ba04c9db9c4c3f08b927fd03982793d6b42

Observation 9c235fba-8e0e-4c66-b221-1238e47e675a · outbound

This paper cites Grok-2 beta release.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grok-2 beta release

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.276520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.460780Z digest=sha256:0c0c97e0bb9979c8ddc75140b8466ab685aef1338f30501954edf9939ae4b62e

Observation 78578b2e-8818-4370-b962-4ba188d86139 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Youtube-vos: Sequence-to-sequence video object segmentation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.265029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.463793Z digest=sha256:89136dc18618fdf6dddce590480e02b5bc3453128f2b0ec7d3adc23f293bcb0b

Observation bed3bb16-6535-4718-8263-7575eca19771 · outbound

This paper cites Qwen2.5 Technical Report.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen2.5 Technical Report

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.466933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.466933Z digest=sha256:394e7a963508fbd88f1ebcae1b70d639d1a6a13a060705f08a5c62cd7040f4e4

Observation a2b826cd-e5d1-4f83-9b2d-45848f42c01e · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.469833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.469833Z digest=sha256:c12ed55c264447212537aa3fd4448e1159a6f4b32d62670f4e2cbf02d752772c

Observation e178e433-1612-4a94-ab5f-f9e3953d7358 · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.472814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.472814Z digest=sha256:ff34cfe5caec26af8d2a03ad2e8d55e2bb14cb5347c56ccd56e92475c7486c8b

Observation 3778f7a8-6891-4d1a-84fd-f2edad6885a4 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.254346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.476387Z digest=sha256:670a6e4f85a3bc4718ec523857943825cf0902a5301a0abbd31ff8440116765d

Observation e68b6a35-fab0-45ca-a137-d16da4d723e5 · outbound

This paper cites Improving 2d feature representations by 3d-aware fine-tuning.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Improving 2d feature representations by 3d-aware fine-tuning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.243822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.479189Z digest=sha256:509a0f641c92f3782e5a3ed94042e47de9ae2398da87e3366dd4bf825f9e5897

Observation 906c1777-421d-4184-888d-adc06cbdecfd · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.481827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.481827Z digest=sha256:da18c42a78eef47836bf33a224e03b8af23d7700dd4611e7b4396663411b058b

Observation e8ab4673-66da-40db-9843-af89c382160e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.485190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.485190Z digest=sha256:1eb919f28e732be15cb10dd7c5923cfceef5e582d6a3fa823e27b754c3dd4449

Observation 7c9db707-929a-4ec3-9346-a4319d2da646 · outbound

This paper cites COMBO: Compositional World Models for Embodied Multi-Agent Cooperation.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models COMBO: Compositional World Models for Embodied Multi-Agent Cooperation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.488149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.488149Z digest=sha256:ae31bcbfed5d00bc5fc1ff885a47d8f22510d1dfbf54944487bbb182754aa98a

Observation ed33a6e8-ec2f-497f-b9c9-448ad37156f5 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, 2024 b.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Llava-next: A strong zero-shot video understanding model, 2024 b

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.231812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.491581Z digest=sha256:699437bdc71698bcf480f216c72c0448848c14be504cb248bed3969a4238be56

Observation bf9ba7c1-8894-43ca-9c45-32e62ae6b963 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.494969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.494969Z digest=sha256:a1315412595d8478ffbe0d66cfdfe21ac8b0d886292c2cbede72246eca2e5fff

Observation eb72600d-c8fd-4187-b371-bee4d690d62b · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:12:39.220174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:12:38.498012Z digest=sha256:c4b1a0f763851cfa9bd7a092b82418ede3edb6a9e83c166e3bbc53f0a0d74cd8

Pith citing papers

Observation 52903efc-fbc8-49af-b197-4bdc8b134722 · inbound

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km cites this paper.

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:37:41.277166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:37:41.277166Z digest=sha256:67e7e8aea2b26d461c5fedf5848e2d582dfac3e50e16076deb60e7bc3341b253

Observation b9d7bddd-8785-4d5c-bb5e-1031eda8cec5 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.883037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:83b0c81b07f06b5df9014f7dea94ea919832bcefb2b12a3624454a9a0e61812a

Observation b428e44c-bccd-4e54-903b-c1e3f9ef82b3 · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.226528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:a051c9653a70fa562720b8859f7c4b62060e017a9111c6f25bd3b64dc10d51d5

Observation 261994e6-167f-4bdc-b274-17dfea95c1c4 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:39.980679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:4d2c26b7c53373c8ba55da01a68b8b70a5527c3136f94fee3552077338d1ffba

Observation febf3a8a-1ac7-4b17-b709-d0e8ce2faf8b · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:fb9d97dd8797bb847407f4d21e8024daadfb212f2c06108d79f068d54e0deaf6

Observation 493af086-78f3-47ac-ba1f-8cd2174e622e · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:52.231964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:52.231964Z digest=sha256:d627d9cf1aad9af6e22d4c780e8b9d3597d56b80262c50d3130b3d99d8ed9285