Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T18:06:09.998507Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2603.16461.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T18:06:09.998507Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-07T13:45:53.346402Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:46:26.794017Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1fcca2d3-6baa-47b7-8752-d99503cc7058 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: European conference on computer vision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588d605b-4154-431c-985a-d2060adfd6fe · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15bb0ec-165c-4d4c-a020-de0815af47f5 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a8376b-b6bb-462f-8fe2-911d0a8d9ffd · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2509.25413 (2025)
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a236d9-50b0-4486-9aa9-2dd88e24b991 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2410.01647 (2024)
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91e9bb2-c90b-47e9-995b-1e675a7f17e5 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2603.00912 (2026)
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d2f36b-d8fd-44fa-af04-709617b99728 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neu- ral Information Processing Systems36, 71862–71873 (2023)
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1e9b0cd-778b-47c5-bfb4-bb6f69abcef3 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 16 J
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54160019-3dae-432b-aaec-93892e3d85f5 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: ECCV (2020)
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8eb9a3c-6c97-408b-94bb-96d489c796a1 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 865af6f7-82c8-426b-b558-9455f2c59ac4 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed21cab-d60b-4396-8cd9-23f130b683c3 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Grounded 3D-LLM with Referent Tokens
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f64a274-e475-47fe-b58a-bd512e9662a3 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.13800 (2025)
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e5a41c-1917-4984-9027-4369647c380a · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1652069e-2d10-4ccb-b348-c4396a15c58b · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF international conference on computer vision
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6760b3f-42a1-4d6a-b51e-d5f753e5cbc7 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc177dd-407d-4269-87e4-f8868b2ef40c · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Generating Context-Aware Natural Answers for Questions in 3D Scenes
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecc57ab-2221-4b00-8e81-684fbf08f0ed · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bde5895-7a6b-4d60-ab7c-f20a72793119 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a134d4a-f2a0-4caa-bd2a-39920d8eadf3 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2601.11442 (2026)
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efb21c4-a27f-4919-9719-3a056ab52254 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Cambridge university press (2003)
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198bf949-7719-4d36-8145-8662a659a0c8 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neural Information Processing Systems36, 20482–20494 (2023)
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4faea79-27b4-400f-8d07-ef075bd40540 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neural Information Processing Systems 37, 113991–114017 (2024)
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98843dd9-ff48-49ea-855b-63b32843dc1f · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models An Embodied Generalist Agent in 3D World
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900369b8-9543-4269-95d4-364584a09cd7 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT-4o System Card
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ae736e-4d64-4c03-82b2-9b1a1fc80846 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d9e8f8-8ab9-47af-bb2e-9b7deed37f30 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models OpenVLA: An Open-Source Vision-Language-Action Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c711c68-0cd5-4ef3-98d1-cbe8c254770f · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unified Semantic Transformer for 3D Scene Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b970326-5b68-49a3-9c94-3448979b1a73 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d1d639-d6e9-4f89-9fd7-6ca76690e7d7 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.22706 (2025)
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddf83bc-5c11-4ff4-8d26-92fcb9e5b908 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: International conference on machine learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2048f2-77fc-4502-be11-51f8fd6b0096 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Depth Anything 3: Recovering the Visual Space from Any Views
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1939b48-2b0a-46e4-ab41-826f6277cc82 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2405.10255 (2024)
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db83715-bbc9-488e-a1a5-06f8bbdab8dc · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Ad- vances in Neural Information Processing Systems (2025)
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23dd2db-0b40-466f-9558-fc66b03e51eb · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Ad- vances in Neural Information Processing Systems37, 23464–23487 (2024)
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b22861a-8f8a-46fd-b74a-89800da0ddc2 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db357836-d2ad-4316-87a1-1dac645ee97f · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee15ea0-69cc-44c5-9b2c-8fcee3ae62f6 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15d34ea-ab75-40ff-903a-ad20c3dda213 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eacb1e50-027d-460a-8a57-898cd3310017 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: International conference on machine learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326b211d-d163-4176-b133-6e05c5826e14 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd1a76a0-11b3-4788-8627-18c2e0d58dcd · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d43fbef-8407-46c1-a296-a98901fcaeef · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2505.23044 (2025)
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6bd746-e000-4fc3-b695-c9b154cf2774 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8811f6d-b96f-4be4-908d-793fcdd8b17a · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2511.18416 (2025)
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad40dd2-a235-4352-8aa0-e4fa2126fb4d · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c85b51-1ce0-4a92-aff2-fab8181a1f78 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0833805c-8384-4d0d-8a36-13af845644f8 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a629d8d-6cb2-4e7c-86c7-512679084b05 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51656720-c56f-4e2d-9d4c-ac62b7aa7b21 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models $\pi^3$: Permutation-Equivariant Visual Geometry Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d5ae86-8032-4c38-8bb0-745e36e66c59 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a8323a6-c07f-4d7a-b729-77691eb92f5a · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2508.11952 (2025)
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d383857f-9b51-417f-b0ad-bd974ab490f3 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79cf545c-addb-4ace-9c0f-ae7939298575 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2511.05491 (2025)
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d77854-2572-4d74-804c-64e05ccdb025 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2601.02281 (2026)
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154fce46-323e-4dc9-9ba1-038e1a5b9367 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2503.22976 (2025)
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e2179d-74d7-4342-9ece-3a69ed43b3fd · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2505.24625 (2025)
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c37f8d-50d4-47e7-84f0-4a726b9ed11a · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046d1603-6434-4f23-b3ef-35111921709d · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.25760 (2025) GAP-MLLM 19
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78606104-cf6a-4120-8cd3-97f27074cf27 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7adcd3-e2ef-4b33-af1b-ab9fc07ef5a2 · outbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models label"and the point’s 3D coordinate in
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d237b7b1-64a0-4574-a2ca-0b54252f643b · inbound
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.