Pith. sign in

Paper Citation Record · LEDGER

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2603.16461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.16461 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:06:09.998507Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T13:45:53.346402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:46:26.794017Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fcca2d3-6baa-47b7-8752-d99503cc7058 · outbound

This paper cites In: European conference on computer vision.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: European conference on computer vision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.807506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.807506Z digest=sha256:8e365f9dbf84a3a0c9b9f549944f4112efdf9d6bfd3c1f73314736ecfe7ece9a

Observation 588d605b-4154-431c-985a-d2060adfd6fe · outbound

This paper cites Qwen3-VL Technical Report.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.811884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.811884Z digest=sha256:27b7472ab575fe1f28c6aa08f7ad48d8b95ca6b534bdde588d15547847df4112

Observation f15bb0ec-165c-4d4c-a020-de0815af47f5 · outbound

This paper cites Qwen2.5-VL Technical Report.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.815651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.815651Z digest=sha256:8a0603509894d17d8bcaebdb5285fbbd4c8cd7c6748f073c8aea80178e6de9ad

Observation 25a8376b-b6bb-462f-8fe2-911d0a8d9ffd · outbound

This paper cites arXiv preprint arXiv:2509.25413 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2509.25413 (2025)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.819064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.819064Z digest=sha256:d6d58c70196efdce1b08ca53fbf292eeda0f3ede701466b3bc63a20eaa75f936

Observation 47a236d9-50b0-4486-9aa9-2dd88e24b991 · outbound

This paper cites arXiv preprint arXiv:2410.01647 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2410.01647 (2024)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.822371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.822371Z digest=sha256:d1d3a85047932ba5e4d3c586df787ab6859d5b34d002b332b50095856532acac

Observation d91e9bb2-c90b-47e9-995b-1e675a7f17e5 · outbound

This paper cites arXiv preprint arXiv:2603.00912 (2026).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2603.00912 (2026)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.825904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.825904Z digest=sha256:4190433c2ca09841a368f4b90ee45a1558df3060c02464610dd272d56b74c308

Observation c9d2f36b-d8fd-44fa-af04-709617b99728 · outbound

This paper cites Advances in Neu- ral Information Processing Systems36, 71862–71873 (2023).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neu- ral Information Processing Systems36, 71862–71873 (2023)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.829387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.829387Z digest=sha256:da9117a899f9f156917c12a30ef3772b1285e8beebcf98504603d589ca9b5bc9

Observation e1e9b0cd-778b-47c5-bfb4-bb6f69abcef3 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 16 J.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 16 J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.832361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.832361Z digest=sha256:8d7988a51afc67e8c0c8d13b6f3196c685bd846ebbb6b39222b48b5bc600ea4d

Observation 54160019-3dae-432b-aaec-93892e3d85f5 · outbound

This paper cites In: ECCV (2020).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: ECCV (2020)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.835431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.835431Z digest=sha256:f19ae2aaa7becfab948f8263bd93143eea8deaf45c93f3ce786e7cf6394b858a

Observation f8eb9a3c-6c97-408b-94bb-96d489c796a1 · outbound

This paper cites an unresolved cited work.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.838718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.838718Z digest=sha256:a7f9a880c129b63747c441b82703728d831cf706aa424225a8d8f6a564dd73e8

Observation 865af6f7-82c8-426b-b558-9455f2c59ac4 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.841836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.841836Z digest=sha256:5905e9e4b1b4135f4359aa94c3287f959c88c4714829a027c3f44a07a5fb618b

Observation bed21cab-d60b-4396-8cd9-23f130b683c3 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.844829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.844829Z digest=sha256:57c5359d7a82160cb729b75e971094f555e2452bce3c5444407c4ffa34f7e698

Observation 1f64a274-e475-47fe-b58a-bd512e9662a3 · outbound

This paper cites arXiv preprint arXiv:2510.13800 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.13800 (2025)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.848516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.848516Z digest=sha256:0a8c156a93d66c0ecc38bc1b03520591c488fd14256fcfba681d953d614bc4db

Observation 65e5a41c-1917-4984-9027-4369647c380a · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.851409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.851409Z digest=sha256:909ffca6ebe8b599680a0af85bc703ee95a6a8529cb3476d66696fca8a061a48

Observation 1652069e-2d10-4ccb-b348-c4396a15c58b · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.854350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.854350Z digest=sha256:f160094de83c37c1fdc65fb68d60170ab1fe30ad41024f64d3c7a36fa1767fe3

Observation e6760b3f-42a1-4d6a-b51e-d5f753e5cbc7 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.857411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.857411Z digest=sha256:1b9fe6794da42edbd6f7ec6dcbc94a5a6ff8d39d8e5c55776b655a4432b33baa

Observation 3dc177dd-407d-4269-87e4-f8868b2ef40c · outbound

This paper cites Generating Context-Aware Natural Answers for Questions in 3D Scenes.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Generating Context-Aware Natural Answers for Questions in 3D Scenes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.860375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.860375Z digest=sha256:bea6af6acfa817af8e27d23a6643b8ad75a17cc1394e0c8a0babfedf811514a3

Observation 2ecc57ab-2221-4b00-8e81-684fbf08f0ed · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.863681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.863681Z digest=sha256:c9db7e48c3b32e811c6f2726f5a25a4a848f752e84510be1624c172ed348de45

Observation 3bde5895-7a6b-4d60-ab7c-f20a72793119 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.866812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.866812Z digest=sha256:4bc4e514de840a886a0069419c6bd75512a67bf62dfef493772ad6a01a88b471

Observation 2a134d4a-f2a0-4caa-bd2a-39920d8eadf3 · outbound

This paper cites arXiv preprint arXiv:2601.11442 (2026).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2601.11442 (2026)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.870167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.870167Z digest=sha256:d09c4345c8bacd420b9abd3dfe435779eb8bfb9fb3cb8aa5cc1ca8cd4ff68780

Observation 3efb21c4-a27f-4919-9719-3a056ab52254 · outbound

This paper cites Cambridge university press (2003).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Cambridge university press (2003)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.872927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.872927Z digest=sha256:1681d817f698f7ac9f932b9faf1aee65afb2cbcf4757328c89e4b83bfb1609e8

Observation 198bf949-7719-4d36-8145-8662a659a0c8 · outbound

This paper cites Advances in Neural Information Processing Systems36, 20482–20494 (2023).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neural Information Processing Systems36, 20482–20494 (2023)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.875787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.875787Z digest=sha256:6a499e4cdf2a88f94ed634bbf2fcf969d72c29dfec64b0f64d7f6abcbc72fc9c

Observation d4faea79-27b4-400f-8d07-ef075bd40540 · outbound

This paper cites Advances in Neural Information Processing Systems 37, 113991–114017 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neural Information Processing Systems 37, 113991–114017 (2024)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.878556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.878556Z digest=sha256:ef123613b46c3a2bcf49ba8addad6de4a239c9f7b4b9a72835d49013d21ce67a

Observation 98843dd9-ff48-49ea-855b-63b32843dc1f · outbound

This paper cites An Embodied Generalist Agent in 3D World.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models An Embodied Generalist Agent in 3D World

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.881685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.881685Z digest=sha256:660870c770a2915642bc3a83be24acf21ef11ed168f6456fd5a0697c086739ae

Observation 900369b8-9543-4269-95d4-364584a09cd7 · outbound

This paper cites GPT-4o System Card.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.884802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.884802Z digest=sha256:4c1c57d9b8d83925a5f5bc83fb6025c9b9a7545eae4bd078040dc522d964f6a1

Observation 23ae736e-4d64-4c03-82b2-9b1a1fc80846 · outbound

This paper cites MapAnything: Universal Feed-Forward Metric 3D Reconstruction.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models MapAnything: Universal Feed-Forward Metric 3D Reconstruction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.887988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.887988Z digest=sha256:12dbc1629c927dbb8e7154704815099998971f229e2bdfb4bb1c0fb7c553f1e8

Observation 66d9e8f8-8ab9-47af-bb2e-9b7deed37f30 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.891673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.891673Z digest=sha256:7eea260f68ee300be40b38823fa5b280ace2b5a96b33e6d74a785f286770a4dd

Observation 7c711c68-0cd5-4ef3-98d1-cbe8c254770f · outbound

This paper cites Unified Semantic Transformer for 3D Scene Understanding.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unified Semantic Transformer for 3D Scene Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.895072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.895072Z digest=sha256:a559915f6c5deecb41a4d9abcc34636afb0cf7d9ab5d9ada65688dd3aabd32bd

Observation 6b970326-5b68-49a3-9c94-3448979b1a73 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.898270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.898270Z digest=sha256:e02fc88dcf8b4136810b303cf9d261b3fe115b180374130950fb78bc2c8b7e4e

Observation 57d1d639-d6e9-4f89-9fd7-6ca76690e7d7 · outbound

This paper cites arXiv preprint arXiv:2510.22706 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.22706 (2025)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.901428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.901428Z digest=sha256:bac65cb417d25d8c1bdf3691fb40e5d0125fb5d23c0c58b7ee1b2ac69696e4fe

Observation 5ddf83bc-5c11-4ff4-8d26-92fcb9e5b908 · outbound

This paper cites In: International conference on machine learning.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: International conference on machine learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.904718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.904718Z digest=sha256:c7085069a47b61e48fd6f550723842aafcf12d1acdd9cd557c802c68d6e22544

Observation 1c2048f2-77fc-4502-be11-51f8fd6b0096 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Depth Anything 3: Recovering the Visual Space from Any Views

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.907725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.907725Z digest=sha256:859ab66a2c2c8066059184ddda54c9df611bc957dafccd7b383cc82e03a63596

Observation d1939b48-2b0a-46e4-ab41-826f6277cc82 · outbound

This paper cites arXiv preprint arXiv:2405.10255 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2405.10255 (2024)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.910861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.910861Z digest=sha256:a1150665feea1a10220b045d498eec63feef33773aaa455445436e9d1c9611a8

Observation 0db83715-bbc9-488e-a1a5-06f8bbdab8dc · outbound

This paper cites In: Ad- vances in Neural Information Processing Systems (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Ad- vances in Neural Information Processing Systems (2025)

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.914100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.914100Z digest=sha256:70ce5ae84ac80a43f73d6b3280ae78a64db3ec19d99746d2c432dc68e75af393

Observation f23dd2db-0b40-466f-9558-fc66b03e51eb · outbound

This paper cites Ad- vances in Neural Information Processing Systems37, 23464–23487 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Ad- vances in Neural Information Processing Systems37, 23464–23487 (2024)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.917331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.917331Z digest=sha256:5791b3a77e4730b12f2d8f33b481a39601cbfb5220278bba4111a3a3fbb34cb4

Observation 9b22861a-8f8a-46fd-b74a-89800da0ddc2 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.920391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.920391Z digest=sha256:09b4365dc2a88b63f0b35e6fec0d21eac274046818620b0f45cfc16ff6fe6c8d

Observation db357836-d2ad-4316-87a1-1dac645ee97f · outbound

This paper cites In: proceedings of the IEEE/CVF International Conference on Computer Vision.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.923398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.923398Z digest=sha256:5f4ae060340f56395a012eb3baa76dcdc9a31698f2c63945cadd0f70ce84f7cf

Observation 0ee15ea0-69cc-44c5-9b2c-8fcee3ae62f6 · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.926358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.926358Z digest=sha256:2bc22e5e1abdc70c66c250d27afe313fb5624330e2c466f5f741553e3024efd2

Observation f15d34ea-ab75-40ff-903a-ad20c3dda213 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.930155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.930155Z digest=sha256:709aba00e2958bd9166db3eb01a552402dbd24457fb1ea0115320fac4d20cb1e

Observation eacb1e50-027d-460a-8a57-898cd3310017 · outbound

This paper cites In: International conference on machine learning.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: International conference on machine learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.933431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.933431Z digest=sha256:5606b79c1c6cc2836b91793a1533087a0f053e0f7a1d6c51699d36c25c061911

Observation 326b211d-d163-4176-b133-6e05c5826e14 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.936478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.936478Z digest=sha256:87c1c221414ff4c36d2c96d5c10e45ce866c6520a346daf06b7f8867cc7b8cf9

Observation fd1a76a0-11b3-4788-8627-18c2e0d58dcd · outbound

This paper cites FastVGGT: Training-Free Acceleration of Visual Geometry Transformer.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models FastVGGT: Training-Free Acceleration of Visual Geometry Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.939650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.939650Z digest=sha256:afac023758ae2f472ae86b9a15ecf3db56a7c27170e1fce3339e9a94a56a9489

Observation 9d43fbef-8407-46c1-a296-a98901fcaeef · outbound

This paper cites arXiv preprint arXiv:2505.23044 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2505.23044 (2025)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.943321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.943321Z digest=sha256:0ebc57ffaa3ee4bc7b012bed6988552a2b27d3f3875bc098805cc1b7cbe6accc

Observation 2f6bd746-e000-4fc3-b695-c9b154cf2774 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.946206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.946206Z digest=sha256:8855cef8ff5537259774e225fb479f6e18a1c634cfe608e1c99f33a6e358930e

Observation e8811f6d-b96f-4be4-908d-793fcdd8b17a · outbound

This paper cites arXiv preprint arXiv:2511.18416 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2511.18416 (2025)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.949484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.949484Z digest=sha256:e9094c65a6d8e34e04216881866366bd8fb9d87d4f17f625855440e646015e03

Observation 0ad40dd2-a235-4352-8aa0-e4fa2126fb4d · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.952498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.952498Z digest=sha256:68deee4f628133d4a1b6512562d6f2c854b310686cc7511be67053c088ffea6e

Observation f2c85b51-1ce0-4a92-aff2-fab8181a1f78 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.955607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.955607Z digest=sha256:fecd2586957f63c4c3dae990669680f7e970ec0443f465840110bf83e7617bde

Observation 0833805c-8384-4d0d-8a36-13af845644f8 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.958696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.958696Z digest=sha256:5ff4fcd94f97d225be7d9502ee667406f870691c5fb3e5516a2169309bd40340

Observation 0a629d8d-6cb2-4e7c-86c7-512679084b05 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.961597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.961597Z digest=sha256:50decc302b60cf35998f68551f5df6967d3b33557c693c47313faf60bf113237

Observation 51656720-c56f-4e2d-9d4c-ac62b7aa7b21 · outbound

This paper cites $\pi^3$: Permutation-Equivariant Visual Geometry Learning.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models $\pi^3$: Permutation-Equivariant Visual Geometry Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.964640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.964640Z digest=sha256:318ddb03c128f0021fd52bd54a2cfc157622f072f31fa191e8f3b2324ac76d07

Observation 68d5ae86-8032-4c38-8bb0-745e36e66c59 · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.967760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.967760Z digest=sha256:df01d57cfb1db1b7038a2a36ff8f487d75eb33046059417c6ee3fc49a4a07808

Observation 0a8323a6-c07f-4d7a-b729-77691eb92f5a · outbound

This paper cites arXiv preprint arXiv:2508.11952 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2508.11952 (2025)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.970802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.970802Z digest=sha256:dfd4e76f81722fca3f6cf12649f4ef5726d4e93085940bb29b7b13c70665ff6b

Observation d383857f-9b51-417f-b0ad-bd974ab490f3 · outbound

This paper cites In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.973850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.973850Z digest=sha256:70d921196da5187dda755c7d80e3917a19bd32e0120b6bf354e280d912eac5e8

Observation 79cf545c-addb-4ace-9c0f-ae7939298575 · outbound

This paper cites arXiv preprint arXiv:2511.05491 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2511.05491 (2025)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.976837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.976837Z digest=sha256:5926946b3bda73f60a9547ea5eef1298ac5cfe2fbc72b0849502b3e53ac73a00

Observation 54d77854-2572-4d74-804c-64e05ccdb025 · outbound

This paper cites arXiv preprint arXiv:2601.02281 (2026).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2601.02281 (2026)

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.979716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.979716Z digest=sha256:7e993b74bb9fabd089ad6860440055806d520c6bbadd7eda054994aa824a27a1

Observation 154fce46-323e-4dc9-9ba1-038e1a5b9367 · outbound

This paper cites arXiv preprint arXiv:2503.22976 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2503.22976 (2025)

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.982621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.982621Z digest=sha256:9a4f2ba7b399d229fe46f31547c6994dc9d05d592ca33f8439117b96a6285143

Observation f0e2179d-74d7-4342-9ece-3a69ed43b3fd · outbound

This paper cites arXiv preprint arXiv:2505.24625 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2505.24625 (2025)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.985621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.985621Z digest=sha256:a5b56c6b82b6dcdc8d0a66cad5387a167236e0c9b1399c2bb2a220c7c73c64a8

Observation f6c37f8d-50d4-47e7-84f0-4a726b9ed11a · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.989175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.989175Z digest=sha256:957f0093356219d7027067a307ca08003cead5207f7559f926dbb96330afb0e5

Observation 046d1603-6434-4f23-b3ef-35111921709d · outbound

This paper cites arXiv preprint arXiv:2510.25760 (2025) GAP-MLLM 19.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.25760 (2025) GAP-MLLM 19

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.992189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.992189Z digest=sha256:218c22b19222efb99601fa4af855ee09585f2de7d99933a77e8ae08ee843504d

Observation 78606104-cf6a-4120-8cd3-97f27074cf27 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.995051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.995051Z digest=sha256:48bed314e6cf69574356e0045cc40fdf7dc8acd6dcc1ef2bc53afa37ef02c53b

Observation 7e7adcd3-e2ef-4b33-af1b-ab9fc07ef5a2 · outbound

This paper cites label"and the point’s 3D coordinate in.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models label"and the point’s 3D coordinate in

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.998507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.998507Z digest=sha256:9e65238f5a476cb93529e77eebb989a11dc058d39fb060bc06465cf13b917545

Pith citing papers

Observation d237b7b1-64a0-4574-a2ca-0b54252f643b · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:21:33.200827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:0cadc80135009550e1609f347ec2cba268f83e33af3ff825da547812a77e55f5