Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:12:38.498012Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 6 inbound Pith citation observations for arXiv:2508.02095.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:12:38.498012Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:37:41.277166Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T05:46:39.977556Z
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2cd9df5c-473e-46db-9a0e-1eefb1f18b86 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cf6149-3466-4b5c-bde3-18d813296273 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03157756-6f0d-434c-bf3d-df149558d8b7 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da41d6c2-eb7e-4e3f-be99-4e02b8ae4d9a · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023a61aa-dea6-453f-911a-67a3005c28a2 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Cosmos World Foundation Model Platform for Physical AI
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 664cab2c-53c6-434f-9e6b-512bfa8ac78f · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Pixtral 12B
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25dd5a9-5c12-403c-9ed8-387fef226c4f · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models System card: Claude opus 4 & claude sonnet 4
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef7effab-fbbc-4fee-8e0a-f5b43ec98a7c · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca597d7-1351-4284-8a0d-39b77dc38e1b · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video generation models as world simulators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0830ce38-d641-4fc2-970c-412f52cebc59 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Language models are few-shot learners
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73dec874-a00c-4e02-83d3-17cd5c77c48a · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatial memory: how egocentric and allocentric combine
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb75277-81a3-4e7c-aff8-73480e991858 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models On learning mechanical laws of motion from video using neural networks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138e8c17-41e3-4474-8df0-a8580409f448 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27a6624-6e91-456d-86ac-9850d9b27fb5 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af65b91-3c46-4259-b5ff-e5b2e8fdbee5 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7255e154-2e35-4724-8474-3cd9e6b63110 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e653a79c-4372-4156-9274-87c14bf4e636 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 216bba94-d7db-4cb4-8695-c53635ec6aad · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spatialrgpt: Grounded spatial reasoning in vision-language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9090c138-07fc-4ccb-b204-1fd92cf0aabe · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4cdeac-312f-41c5-b891-9a6a14f63b6f · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a1a71d-81a4-43bc-a62d-bbe78f149e65 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea616791-2452-45a3-978a-2594d6df6fee · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Myers, and Anna C
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41f1e65-2363-4bad-ac67-c0680f910318 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc3a3a9-09a8-4e53-8863-7f09d3716f67 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378e8e72-bcfb-4c50-ac30-624e52751d48 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Palm-e: An embodied multimodal language model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4be8d57-ab48-46e1-818a-b13203b91d37 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92559efe-a876-4d36-977f-ad9c8d05ab43 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Large spatial model: End-to-end unposed images to semantic 3d
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5e0ebd-133a-4469-b99f-ee3463fcf341 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6d57b7-444e-4382-91a7-24bdb6bbe574 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Freyd and Ronald A
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6c2ad2-e003-4f25-b766-b51bcd6c89cb · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78cae08-3807-4eea-a7d4-d3b6817c0802 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c873d276-6d9c-4465-aca0-e444a94ae04b · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76488af5-b9c9-4fc5-b18a-ea10b773592d · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Ego4d: Around the world in 3,000 hours of egocentric video
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c6eb71-8bda-484c-a9a4-72e10f1ef3d5 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768265ba-529a-4ec3-a474-af06077ec43b · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mojito: Motion Trajectory and Intensity Control for Video Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e1ad44-514d-4fd2-acfb-2109865dbaaa · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models GPT-4o System Card
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91eb9c84-dedb-47fd-a964-c26841f2dfdd · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visual perception of biological motion and a model for its analysis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3926e618-9310-446e-92b2-7a6c75d3ddf5 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8aa67fb-5870-41f0-be5b-d5ec4fbcb5f7 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models OpenVLA: An Open-Source Vision-Language-Action Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c06351-8938-4948-8fa7-03c4a5aa02e3 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Decomposing nerf for editing via feature field distillation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877b0cd2-1373-43a1-a3ba-72745c29571c · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models On space-time interest points
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef2fe79-e9c0-42e9-b806-8a54f417adbd · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Denker, Donnie Henderson, Richard E
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19da999-8d1f-4410-81c2-e18f0e18ef1a · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 431ecfc8-b97f-4732-9738-36b17afad5ce · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Seed-bench: Benchmarking multimodal large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4556114-2f60-44f3-887f-a6f9b9ad2e81 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967cacbc-4c71-454c-96d2-8db4b59de9d3 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ee9794-d3cf-40dd-a837-5ce11f4ea952 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoChat: Chat-Centric Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f45c45fc-5204-4e98-a655-3cf0b5ec2dc3 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mvbench: A comprehensive multi-modal video understanding benchmark, 2023 b
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c106f83-17d9-4aa5-a423-d32b6d1a2bdd · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9385a981-cc8d-474b-95bd-8d66dcdb2bd6 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models 4k4 DG en: Panoramic 4d generation at 4k resolution
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 018d2dbe-f6e4-4b6e-8b39-75981ee43f78 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed4d6db-0496-461b-b66f-3ea999ddd1be · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Visual instruction tuning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea025cc-1422-4aa1-9fc7-c6bc4b033163 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models World model on million-length video and language with ringattention
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51dc2eca-82f2-4fda-9721-8dc7dd0f5e69 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916677d5-7baf-4fd4-8920-2e4d2d5c9dfd · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2891cc0-810f-47de-a0ca-3fdf9bde578e · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec286d9c-b5c3-4e5f-8d79-9cf591fb5f30 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e62240e-473e-426c-a6cb-fcad01a37c68 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8ee27b-effc-4a6c-a1a1-ab9a57d7f7fd · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Marr and S
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ea2a558-96e3-45a5-b926-575bde0fd5f7 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67684500-24fc-4457-a575-c13145b72f96 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b720db-3f87-47e2-b9ef-84a16d44b79a · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Sceneteller: Language-to-3d scene generation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00ff1076-2416-46ed-9b02-321ced9abc42 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Hello gpt-4o
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16a0106c-88aa-4b60-a40b-59df9e41021c · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb64b7cd-bf99-476e-b198-b267123171ca · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models A benchmark dataset and evaluation methodology for video object segmentation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fc0dd45-2e18-43e2-b8d8-060dc2ead32d · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models The 2017 DAVIS Challenge on Video Object Segmentation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1491ac1c-17b6-4d41-80bb-7c8c12409df4 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Improving language understanding by generative pre-training
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f338b6-8a45-45b4-b407-6057abf106f6 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Learning transferable visual models from natural language supervision
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c004b0-1cea-4029-9754-68e1f4137c20 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Learning to localize objects improves spatial reasoning in visual-llms
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b53dbe3-ed92-42eb-ae97-181e0f89d6b8 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Two-stream convolutional networks for action recognition in videos
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5d69fda-8888-4ce4-8f7e-f77463128fea · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Spelke and Katherine D
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea651a9d-1b27-4bff-8db3-46c22c15aa32 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a2386b2-d0ec-4ed5-ba0b-9cc7e2f16731 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9915e8-04f5-4ca1-8b26-eb2b4cee5e84 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9945a242-91fa-48b0-b0ef-6ae9974d16ea · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Wan: Open and Advanced Large-Scale Video Generative Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464dbc8e-7267-4361-8499-dae2e896ccb3 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c0ae7d-3c59-4167-ae14-2db5dc8abad1 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Action recognition by dense trajectories
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13f125b0-b270-4d4e-b194-a727aecc27fb · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 018498f3-03d9-4b86-b582-08f36c393dc4 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95d03c2e-bec6-456a-905e-5231ad6d9159 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f667929-c395-473a-ade8-bd8a3d499292 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Internvideo2: Scaling foundation models for multimodal video understanding
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e018751-258d-4358-8242-05dba5a50bca · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5650af5a-e4d5-43aa-8037-9f400c9a7c0e · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Finetuned Language Models Are Zero-Shot Learners
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63d73dd-d739-4c97-b02d-66b3c0a5b79b · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Chain-of-thought prompting elicits reasoning in large language models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e3e534-fd64-4804-86fc-5ea62873d6f6 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Cat4d: Create anything in 4d with multi-view video diffusion models
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e0bbea3-32fc-46ed-97e6-3782efc20926 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c235fba-8e0e-4c66-b221-1238e47e675a · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Grok-2 beta release
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78578b2e-8818-4370-b962-4ba188d86139 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Youtube-vos: Sequence-to-sequence video object segmentation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bed3bb16-6535-4718-8263-7575eca19771 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Qwen2.5 Technical Report
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b826cd-e5d1-4f83-9b2d-45848f42c01e · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e178e433-1612-4a94-ab5f-f9e3953d7358 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3778f7a8-6891-4d1a-84fd-f2edad6885a4 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e68b6a35-fab0-45ca-a137-d16da4d723e5 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Improving 2d feature representations by 3d-aware fine-tuning
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 906c1777-421d-4184-888d-adc06cbdecfd · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ab4673-66da-40db-9843-af89c382160e · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9db707-929a-4ec3-9346-a4319d2da646 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed33a6e8-ec2f-497f-b9c9-448ad37156f5 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Llava-next: A strong zero-shot video understanding model, 2024 b
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf9ba7c1-8894-43ca-9c45-32e62ae6b963 · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb72600d-c8fd-4187-b371-bee4d690d62b · outbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Llamafactory: Unified efficient fine-tuning of 100+ language models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52903efc-fbc8-49af-b197-4bdc8b134722 · inbound
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d7bddd-8785-4d5c-bb5e-1031eda8cec5 · inbound
World Simulation with Video Foundation Models for Physical AI VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b428e44c-bccd-4e54-903b-c1e3f9ef82b3 · inbound
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 261994e6-167f-4bdc-b274-17dfea95c1c4 · inbound
The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation febf3a8a-1ac7-4b17-b709-d0e8ce2faf8b · inbound
The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 493af086-78f3-47ac-ba1f-8cd2174e622e · inbound
The TIME Machine: On The Power of Motion for Efficient Perception VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.