Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:26:19.763700Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 12 inbound Pith citation observations for arXiv:2412.00493.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:26:19.763700Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:11.071380Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T01:00:51.358037Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8877d8d5-9b7b-4c2f-a68c-b79cd714b710 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f390bf75-f74b-47ff-9b99-688b48be521b · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scanqa: 3d question answering for spatial scene understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7fcd4a09-f3fc-4b44-ab77-dce2bc4a78b8 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3djcg: A unified framework for joint dense caption- ing and visual grounding on 3d point clouds
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e685faa3-2175-4ab5-b915-ce2ff4f6bee9 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Guibas, and Fei Xia
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b0c42af2-40ad-4956-a903-0094a211f815 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chang, and Matthias Nießner
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b3d87e49-2dad-4b71-8073-7582904abb67 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c2e5edaa-2613-4029-82c5-9f74a70e139f · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bebb84cf-978f-4307-b76a-9a3283f00fc2 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Language conditioned spatial relation reasoning for 3d object grounding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3d27d9f4-8102-4ca0-bb97-b9b222df0427 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding LL3DA: visual interactive instruction tuning for omni-3d un- derstanding, reasoning, and planning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64981b8e-d522-4f13-bfcb-628ef9b5f471 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Grounded 3D-LLM with Referent Tokens
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a73b7cd-cad0-4227-b25a-ed94845beb48 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Intern VL: scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f7b8c156-85f8-49e1-8e0d-a9c16700aee4 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chang, Manolis Savva, Maciej Hal- ber, Thomas A
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8f6c05eb-9542-444b-88b4-bca6494f8448 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding An image is worth 16x16 words: Transformers for image recognition at scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cec3e35-4029-4ede-9680-12b6169a77db · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9650f3-8338-44f2-95a8-698d8cf91340 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ff64a6-c90a-4a73-9ece-323b7e22b0a3 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3d-llm: In- jecting the 3d world into large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f82e33e3-62e5-412a-8276-a6204f0f24b7 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599d0e1f-4988-4d43-b4fa-89a4aa9731b5 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding An embodied generalist agent in 3d world
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b9a8c148-edc9-43d7-a551-e7be0aadae39 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Multi- view transformer for 3d visual grounding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12b43e78-d485-486a-899d-56e04ce7f700 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d9ba2e-890f-4984-92e0-21f24b917066 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding The budgeted maximum coverage problem.Information Process- ing Letters, 70(1):39–45, 1999
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c3da25e1-c635-4943-889b-e8c17fdaddf3 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding OpenVLA: An Open-Source Vision-Language-Action Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90534326-cfc1-4102-81f9-e41489960c3e · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e9609b-219f-4beb-8668-26cc1a514994 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c9a8f55b-b79c-4c41-ace9-3cb9e484c3e3 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Mvbench: A comprehensive multi- modal video understanding benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c14c733d-09e3-42b6-83dd-9547ef8d1f6a · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding VILA: on pre-training for vi- sual language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a12a8f55-5848-4ce5-998d-15c55039cc4e · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Multi- modal situated reasoning in 3d scenes
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79d48281-b306-48e1-ab6a-3d2d66438a15 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d599c639-00b2-4896-b576-4dc0849cfe42 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Visual instruction tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e9ab469b-1570-4b36-a5ef-5d045d1d3b1d · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61d4049-8934-4e47-beb4-c54b5f41605a · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding SQA3D: situ- ated question answering in 3d scenes
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3f945e42-0fe6-46fa-9aa8-0eb9c9519b90 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Openeqa: Embodied question answering in the era of foundation models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5b47376-fa21-4202-8631-d992952119f1 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding GPT-4 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086660a2-e875-4904-a7e4-89324de6cea1 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bac8b147-e9c3-4e62-885b-61a83f3a8f49 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Representation Learning with Contrastive Predictive Coding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b54314-aae9-4b88-abc0-04277f916161 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cff1c11f-7e9b-4164-81ff-c6a424e112c9 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Lawrence Zitnick, and Devi Parikh
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fbce2f75-bd48-45c4-8f80-f34edc7488c7 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding RIO: 3d object instance re-localization in changing indoor environments
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c65460b7-88a2-436c-bc0e-8cb8f8c5e46b · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db3730d-6d90-4267-8779-d8710368023b · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ccbae5-c6fc-403c-bf7d-f4681a58639c · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unleashing large-scale video generative pre-training for visual robot manipulation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 08991898-4268-4499-a3a7-b6eaba3c4f16 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Pointllm: Empowering large language models to understand point clouds
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dbc7b6b9-57b6-4383-98cd-9f99f5584892 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Qwen2 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c9e5d1-6395-421b-a21b-7d656022ced1 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scannet++: A high-fidelity dataset of 3d indoor scenes
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 861e3871-867e-4a4b-9e5c-337351f765a7 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0db287ca-307d-4243-8995-3398c5513ad7 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 18522524-38f7-4c96-bc66-9c1e775d4560 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Video instruction tuning with synthetic data, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 62b66678-86f0-4874-8129-8420e6eaa992 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3dvg- transformer: Relation modeling for visual grounding on point clouds
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 39cc2573-ba74-45a5-9013-3404aef4e92c · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Towards learning a generalist model for embodied navigation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 676546bc-40b0-4bcd-bcc4-bb54fa735bdb · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scanreason: Empowering 3d visual grounding with reasoning capabilities
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 88242134-1269-491b-a8a9-efc1291d4763 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a56442-6815-442b-b74e-271e088112a5 · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f09a9d05-930a-4648-a5fa-651c141dc30f · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unifying 3d vision-language understanding via prompt- able queries
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 14b9ca5b-87fa-4197-9d54-6a645cba941e · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding sos” and “eos
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 02f8bf64-0933-4f8f-842b-ff07a356618b · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding ZT” denotes zero- target, “ST
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a89ac1e7-e156-434d-a51d-0dd70c3ce81e · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f3ac7b5b-73ec-45fb-ab0a-464dc426e83d · outbound
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 1, 2, 5, 6, 13, 14
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0c04a790-97a6-474a-b317-50b3f1a42402 · inbound
The Internet of Large Language Models: An Orchestration Framework for LLM Training and Knowledge Exchange Toward Artificial General Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950c6378-f069-49eb-9012-56ffe44cb7b6 · inbound
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f1dc900-6b44-4a4c-97c7-1fb7e79d48b9 · inbound
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31acac97-89dd-432e-855f-aef3ea07a7bd · inbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c60e752-4b9b-46f1-8d89-b0217515b7c5 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a4c06c78-5058-4645-811b-0778b0d90f73 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · inbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0cd3b5-f8a4-4fb0-95d8-c5b13fa3842e · inbound
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f03306a-48ba-4fb2-813d-0922a73fea03 · inbound
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6f6f27bb-e7d2-4e8a-8ae7-ca3dc3495693 · inbound
A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40688cb6-5a6f-412d-afde-7fe87b94e9ca · inbound
Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34fcf9c-3f7f-4b4b-b317-abf2803f0958 · inbound
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.