Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T09:09:24.905094Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 0 inbound Pith citation observations for arXiv:2607.20868.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T09:09:24.905094Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8125c7ac-7cdf-4e38-9344-8fc03592cdb2 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Anthropic Model System Cards,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b3c9d1-5eac-4fcc-b675-245d1353e068 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c603d0e4-fa8f-48d3-91b0-3c234ede0cbd · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dfb85d0-7dcb-4255-b922-e76af0dfe948 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a52870-9495-4760-80a0-20d6d78dfc39 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MMCode: Bench- marking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f30c3b6-0f4a-439d-ae62-4ba10ff86033 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e42d47d-a9aa-425d-bbe5-8ce38f1520ea · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Emu3: Next-Token Prediction is All You Need
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a587680-5cd6-4e9a-ab65-1ac8f160b1f8 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? NExT-GPT: Any-to-Any Multimodal LLM,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b646d324-077b-4ef7-bf9e-3076e5feab10 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MiniGPT-5: Interleaved Vision- and-Language Generation via Generative V okens,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e092185f-4012-463c-8ad9-5caf242f27f1 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Introducing GPT-5,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5414616-3fa1-4111-81e7-c2254d124fb8 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a219f89-569b-4bcb-983f-01c58090da0d · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Gemini Achieves Gold-Medal Level at the Interna- tional Collegiate Programming Contest World Finals,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 999c13da-e07a-4e30-886b-2152ef9bb75c · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LMDrive: Closed-Loop End-to-End Driving with Large Language Models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ca3259-f664-4575-9707-708e924734c5 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? DriveLM: Driving with Graph Visual Question Answering,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875c0729-2bc9-4309-8560-39c395827fa0 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff14f369-bb9d-4d3e-a525-eefa7817b933 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? OpenVLA: An Open-Source Vision-Language-Action Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1097f00-afa0-4748-85b3-b4a86d71b181 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932ca2e5-c28f-44e2-9e9a-d088986d966d · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0694721-20aa-4bcc-82a2-4ce2beb7384d · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? PaLM-E: An Embodied Multimodal Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ae75f9-3ea0-4dc3-96ab-dfaa6c831990 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50629123-22f8-4624-bc15-ab6f08732a88 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MMSI-Video-Bench: A Holistic Benchmark for Video- Based Spatial Intelligence,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f47daa6-4eac-4045-8d67-766dd6457b75 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2a0d0c-1278-4573-9d1d-47c318a0ba71 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681bffab-3e86-480f-b187-6841f26bc275 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Cambrian-S: Towards Spatial Supersensing in Video,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbff37c7-3f65-47ca-ad89-4b8ee1f4dc27 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46b15c0-31fe-4bbf-8168-4f84d3b93821 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0fc3fc2-9ee6-46d0-b1ba-665cc9a35428 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? STI- Bench: Are MLLMs Ready for Precise Spatial-Temporal World Under- standing?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d91ab7-bc77-41c4-9f5b-7740f17bfe69 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? OST- Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e438968-41c0-4e76-87c2-3143e44f20e1 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fcb42e-2cd8-47a6-8527-7d5036bf4f43 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 632c029d-ae40-4b0a-9287-5be632dd2d34 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? DSI-Bench: A Benchmark for Dynamic Spatial Intelligence,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f25c53d-8930-4d28-90c6-36fd72a12aaa · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a937fd-b478-45b6-9168-acefdddd79b4 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VLM4D: Towards Spatiotemporal Awareness in Vision Language Models,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c921455-a00c-4eda-9ec9-2b29bf2ca76c · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Ego4D: Around the World in 3,000 Hours of Egocentric Video,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a9cd58-b35e-4787-b9ea-d3a24d45bd36 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbee15b2-dffa-44e3-8b21-9ed07d5187f3 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? ScanNet++: A High- Fidelity Dataset of 3D Indoor Scenes,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc901f7-7838-409d-9765-20db563219af · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ced9224-8fc2-47cb-ac81-07854632bd49 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Scalability in Perception for Autonomous Driving: Waymo Open Dataset,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bb028a-ca57-4dee-b0d1-c42c7325a144 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f9d890-4491-4575-89c3-834484b10ef8 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? OpenAI GPT-5 System Card
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247ccc9a-97cb-43c8-9e23-f0a745cc11c3 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Google DeepMind Model Cards,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f70b11-33f6-48ac-a5db-a7d93346b331 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337004c1-8c65-4208-9792-2559028c0d95 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Xiaomi MiMo-V2.5: A Leap in Agency and Multimodality,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d380a1-a1f7-4aa6-9722-3a6bf90d670d · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf4b523-2e5a-4b0e-936b-6018b2bd5414 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Qwen3.5: Towards Native Multimodal Agents,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6455fee5-e63b-46c8-9063-9e292c0667d1 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da3ae93-117f-48fa-aaf2-e6e92ac16acb · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? GLM-4.6V: Open Source Multimodal Models with Native Multi- modal Tool Use,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e90f630-c0bb-4991-b1e6-85310d822e28 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77c7cc0-872d-45d2-ba58-5634e812fa33 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfc4348-ae37-4563-81a8-274ba30d4094 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Spatial-SSRL: Enhancing Spatial Understanding via Self- Supervised Reinforcement Learning,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84a6140-9ecb-4bcc-9046-4edac70496b4 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03be65a3-aa7f-4204-8ec4-665c49ad8588 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? SEED-Bench: Benchmarking Multimodal Large Language Models,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89616508-088f-4ae6-90f7-2d7ae391846a · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MMBench: Is Your Multi-modal Model an All- around Player?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6948210e-ade8-4669-bdde-8e30be83d835 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e166ab93-87c6-4c3b-b29e-a20fccb9a6a4 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d929df-7a87-4e2f-be7f-833a94d8cddd · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb8d6b51-583c-4c71-aee1-e45228e295bc · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d05bc6-eccd-48a7-b2a8-d7b69db06a4e · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41dc602-b913-4dbb-9eb1-9da3df568b13 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VI- TATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524bb6cc-c8cd-42aa-b39e-26b785af56bb · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Multi-modal Situated Reasoning in 3D Scenes,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd3159b-4ab0-4644-8bd1-b6ce4ab685b2 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? TempCompass: Do Video LLMs Really Understand Videos?
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d18da41-9fde-47ab-9a22-e8f88c443101 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? OpenEQA: Embodied Question Answering in the Era of Foundation Models,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde51b8d-79f0-4ace-81d5-51ad129d2d86 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8dde046-2a09-4779-9adc-afc15e2487f4 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a345f21-38be-4288-af8b-a5b1c96cb446 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? TOMATO: Assessing Visual Temporal Reasoning Capabili- ties in Multimodal Foundation Models,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9858ce-15bf-4eb1-aa57-f8cfb0b83e54 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for 14 Vision Language Models,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6551dddb-c522-45a8-9185-daaba9d91f10 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de799c9-32c9-4146-a16d-4d1b0ec32444 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39dce1a1-745c-4a50-b46d-6c43222b2287 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaVA-OneVision: Easy Visual Task Transfer
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882ebf22-81cb-44cd-b186-957cea2eff66 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VILA: On Pre-training for Visual Language Models,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 883af380-5307-4f7e-ab65-3a07c3726633 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Long Context Transfer from Language to Vision
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a87c986-facc-45bf-96c2-e2730d1a1934 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42219e3-3d23-455c-a964-dfce1e82254e · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57eb1c8-c706-43c6-9117-8cb5704b1a7c · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e38d04a-b57f-4306-9cc8-c17f46773ee0 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743219f6-c244-41e2-8a38-3154353896ba · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? SpatialLadder: Progressive Training for Spatial Rea- soning in Vision-Language Models,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc2f67e-6117-4bc4-8b82-5fb77958e9ed · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b0b770-7908-4719-8dff-ecf4885d6455 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation,
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e93bd3a3-2d60-4ff0-b19f-c9060cbf3e57 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Visual Spatial Tuning,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33160b22-2cbc-44e1-8997-73fbb7384a94 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Make Geometry Matter for Spatial Reasoning,
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e848a5f-55d4-489b-aa0f-50fabea23b58 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Think3D: Thinking with Space for Spatial Reasoning,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd70c46-7875-4f27-8e60-647ad55a4dc6 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580c64bb-00ed-4e6d-ab08-d8e487189a6d · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaV A-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding,
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0b883b-ef1c-4982-a3ad-26960e6d9e7e · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280bde41-c338-4562-a0c0-b2dcfa96dfbf · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Intern-S1: A Scientific Multimodal Foundation Model,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a91b1af-aaaa-408e-9f59-3c47bbe65175 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Intern-S1-Pro: Sci- entific Multimodal Foundation Model at Trillion Scale,
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c89f018-dcca-4964-a10f-5587d10da16c · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Large Language Models are Zero-Shot Reasoners,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587cae99-1559-4369-96cb-9b89e10aceaa · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c507d2f-5628-4032-8526-de369a708b92 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Plan-and-Solve Prompting: Improving Zero-Shot Chain-of- Thought Reasoning by Large Language Models,
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b23c6ca-48b5-4ef3-9cc8-b2fafe368bfb · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VGGT-$\Omega$
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94e877c-cabd-403b-8e2a-0c82c5592813 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? W AFT: Warping-Alone Field Transforms for Optical Flow,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598cb3f3-053b-4e6d-be57-2e196cb507a1 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Don’t Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a120e0d9-4958-4b31-aeb2-b0ab30f7dd23 · outbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? Expressive Body Capture: 3D Hands, Face, and Body From a Single Image,
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.