Pith. sign in

Paper Citation Record · LEDGER

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

As of 14 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2608.10864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10864 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:33:29.808181Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54a1a48e-541c-4a26-99a9-ac63c5153957 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.118241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.118241Z digest=sha256:ef89c765a669a4cf9d4487314ba958799c6b76231cc6afe4d363f277863c5d3b

Observation 1f84e78f-dfcc-4e6e-8447-694e34bbad57 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.168523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.168523Z digest=sha256:0bd870fd91611350b27559fb858a1ac8966701899f954189fa9770e9288938c6

Observation 72786601-5a8c-4972-bbd3-73db14d1f467 · outbound

This paper cites CUP Archive, 1967.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models CUP Archive, 1967

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.208043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.208043Z digest=sha256:4b0fc326b2027c1f98fa9340a9620fc5c7530dd9163f2a5b2b2bfa98297ae2e3

Observation bb9cd8f1-1932-4808-9c04-adda660b79d2 · outbound

This paper cites Henry Holt and Co., Inc., New York, NY , USA, 1982.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Henry Holt and Co., Inc., New York, NY , USA, 1982

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:32.032179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:27.258250Z digest=sha256:33703baf95aab23d5c30ac04337849dad26af6d9636abcbbde767e60e584701c

Observation 8b75e6f2-0426-4b6a-847d-ff22eaf69d19 · outbound

This paper cites Number 6.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Number 6

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.296641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.296641Z digest=sha256:47bc1a2fea37a28513a0841ad073b8c5b0a9b4ef25626d99b2068c271b5112a7

Observation 9c1485a0-d1ed-4f5d-898d-49463389cea6 · outbound

This paper cites Separate visual pathways for perception and action.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Separate visual pathways for perception and action

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.334241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.334241Z digest=sha256:d7620ecc121182cee88415ced62f9f9efe3aac9ee55b048bd37bf4e57b8664a3

Observation c70e74d4-ee1f-4bb0-a477-e732b8540558 · outbound

This paper cites Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.374416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.374416Z digest=sha256:c1955f4e66b69bd6bc89d423ed5e5f2a555becefcc68942b8f4894aa2f77dd74

Observation 53fa584c-0db6-46e2-a28b-fd03061b08d9 · outbound

This paper cites Oxford university press, 1978.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Oxford university press, 1978

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.403655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.403655Z digest=sha256:e9d0a7e095028e24f0a8eb5b23d2dbfe3f7da18ccba4ee20a31384e4e663811f

Observation 3d47e3f0-36cc-4b90-b178-965802fae45f · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.434416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.434416Z digest=sha256:11e33abc21636f2140ae1727f4afb5c56a7e824186600bb1870cfbb0c084dcea

Observation 04ee055b-8d94-4616-9a27-fa3eac91f6a3 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.467684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.467684Z digest=sha256:33f76d62b34b288a6054322f3d45b374173cb3c96b70de0f6bd8bd52c5842ca2

Observation c274fcc8-ac9b-4312-b3d0-4d118e2a2623 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.488894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.488894Z digest=sha256:9599b85fa17d2f778acfdb667cbaa7df87e710ea67f163af234c7ac9efdf95d1

Observation f4020c15-82c7-4b85-b0ef-3ff96bf5ab12 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.513243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.513243Z digest=sha256:a8725653a3c35dd9fda25e874509c1b2237002a6b6846b7f12099a65249e1f45

Observation ecd75291-ac67-40ae-9e16-39afc05ac79f · outbound

This paper cites Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.536117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.536117Z digest=sha256:e0b21240fc24fccc42f97afa7ac4ba571b17607a8e5c0337aa3278949f84a0d2

Observation 0f5fb8ed-eb7f-4d78-a561-d48b5f067980 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.559154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.559154Z digest=sha256:01f623ef2a0619aed2ce087db241fd643c10e13ba0717b2d25d98e5366cb2f19

Observation 205a6793-af32-450d-9f26-0bdfd8316df2 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.597794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.597794Z digest=sha256:017810bb784e6c25d3fc9030fa656b316df91a1077496942ff93aed1f76ee31e

Observation 61b7547e-1d9b-4a8b-9164-36bb4225850c · outbound

This paper cites LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.861793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:27.627853Z digest=sha256:42a2eec4187c8a32aaed70a986dfe14c95e7a6b02f46de8b9c46f90b0f14c05b

Observation 7bb53d38-73a0-4035-a903-7194cffba8ed · outbound

This paper cites Probing the 3d awareness of visual foundation models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Probing the 3d awareness of visual foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.851699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:27.662301Z digest=sha256:9154d892eeef7799f17b7b28baa2456157562d3c667a1fbee9206187a6b2fb0f

Observation 5cf0e01c-42fd-495e-a90d-a30d3d2c7c98 · outbound

This paper cites Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.693368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.693368Z digest=sha256:4daf5150660cbcdc25f177f6b6a5ac2d5ce31c3f3fc71f7f9b4a2d6957b26b86

Observation 87d6ce1f-049d-4151-902c-6f4b92e319bd · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.750364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.750364Z digest=sha256:fc78f9afa4f3bedb598ff66212c50f557dc20f4b94d36767f44bc810554d47b6

Observation ccddb12c-aba2-4058-990e-f3875f62cebe · outbound

This paper cites Vlm4d: Towards spatiotem- poral awareness in vision language models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vlm4d: Towards spatiotem- poral awareness in vision language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.802712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.802712Z digest=sha256:7a2b175de3e772941e3877cd379b9cd43309cced04f2e4d39be07f58d68960c9

Observation 4a06d401-c430-4d9c-bbd9-3f0656262306 · outbound

This paper cites Cambrian-s: Towards spatial supersensing in video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-s: Towards spatial supersensing in video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.915456Z digest=sha256:9fdefffb9121ddce2fbb5bb99a91bf2364c79b194d2cf312979184f85af2d9ea

Observation beac19f4-b222-40b6-9e7e-69cf97e0fcbe · outbound

This paper cites Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.973683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.973683Z digest=sha256:af57c88072d789ee3ceff145868ea269556213617209b54096a5ee86bdbb80cc

Observation 4647a97b-0e1f-4c31-bcbd-a85f98f5f94e · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.004225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.004225Z digest=sha256:641ef772bee402883c80204bec10986693a7c3db4c6e901d42be4c6303ecac3a

Observation f434eff9-3b11-4021-8934-ecd778e162be · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Distilling the Knowledge in a Neural Network

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.057213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.057213Z digest=sha256:b3e41f5d6d3cc26d0db9e1c9f66b6c5d81dcde1594b5221ea8b4087e8b8c2da5

Observation c976a4dc-f4a9-40f3-b939-dd1c9a60943b · outbound

This paper cites Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.803965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.107936Z digest=sha256:65570a4753cc58b8a0a6a9be3372b7fa157f9b99110423d5854419906f42befe

Observation a77fa326-3b6b-490e-bc6b-4cf2003d5c93 · outbound

This paper cites Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.688307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.143957Z digest=sha256:42297f54a09f1ac1674e0a7443aecc42883727a7d154e5625fa252a589eeac05

Observation f1a1f640-c5ea-4332-b960-e17a12b4f2f1 · outbound

This paper cites The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.182390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.182390Z digest=sha256:b44001f25db4dc2cd7f7cf06a81be1794c96d575073ce506161421f2cbe95906

Observation 0b40abdd-a5db-4460-8a8c-2ef393025aae · outbound

This paper cites Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.212915Z digest=sha256:87b080357e6e51464f23e47bbd91f543c8603ddad607f53bfeb46311c6ecf6d3

Observation a6faf50b-53c7-4948-b050-b6ca50e69b8f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.247419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.247419Z digest=sha256:9890fa5ab6c7a9413c95214c66dbbf2a3560d76fb3693f66532743581f817ab9

Observation 67da9146-8180-461b-920c-a09a20989136 · outbound

This paper cites GPT-4o System Card.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models GPT-4o System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.289469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.289469Z digest=sha256:d91b87772bf4d5aed63faab6a030cc382cfbf6217a8666a86e6676b07a1fe72d

Observation 7c938b5e-59a1-4a40-bc81-2a4b0f23cbd7 · outbound

This paper cites STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.317477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.317477Z digest=sha256:5649abe58fd65d99a5670e5350efbc564a1a4c340e4974cb1cbd52125480cc0c

Observation d9ce87d5-a09d-4627-9a8c-8812128504eb · outbound

This paper cites VLM4D: Towards Spatiotemporal Awareness in Vision Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.342219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.342219Z digest=sha256:354f13e298b969c6e539c74928e5700ca2bd43ef7133bca190e4c2bd4fcdbf24

Observation 6ccf82cd-9f78-4f08-9861-33308bddca74 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.364382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.364382Z digest=sha256:66afd5c6a40898f1e5573431dee8c7d59982da4747e2da0c092b5a8a621b3bae

Observation 1c083e24-1267-4b2f-94ec-a3fe365ffef1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sigmoid Loss for Language Image Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.381618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.381618Z digest=sha256:47a7de9751bd41cfc32693575a98c05284706b14306c80b92a43199ada9f6da3

Observation 8e3a871a-0a6a-473f-a70b-2be1a1d4fa0f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.399065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.399065Z digest=sha256:ebc12c652e2a20317dda8fe68d9257595559965ed308887702fd6785e78f450e

Observation 96ab4e6c-53d2-426a-abc4-f0e554584a1f · outbound

This paper cites O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.460314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.460314Z digest=sha256:d469298d68bc334922f1084334e33b6c0289ff404696b105cd3bb075fa36947a

Observation f65e9dfb-4032-429b-9eef-45796e4626c8 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.508799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.508799Z digest=sha256:25c02faced936d2e7b5e79ac6f2a9c0646a24f2bb069ec6c3992771fdf1ab81b

Observation ebcc31ef-8308-4671-a011-7d8922eaf6e2 · outbound

This paper cites Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.562622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.562622Z digest=sha256:77afa48012611fd1a550934345f5dff20098b45797f3bb03f20097b476bc7d42

Observation 06b1887d-f586-4ca0-8901-29e9cdbf6dd2 · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-S: Towards Spatial Supersensing in Video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.585047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.585047Z digest=sha256:a28a332456999e05fbbcd3f15bf1656e6caebd136cd94bf359c3584de3ba9ff8

Observation da5f7c1d-cf31-4b45-84aa-571d67dcbbf6 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vggt: Visual geometry grounded transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.614240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.614240Z digest=sha256:5c7b77dc27ca31a822774600806fe7d843d8abd606a008eef6e8998cb44c6eb4

Observation 10bccea0-33e6-4810-9190-2e9cd6e583d6 · outbound

This paper cites Continuous 3d perception model with persistent state.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Continuous 3d perception model with persistent state

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.541921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.641131Z digest=sha256:c52292c74d01f08a04dc992d3af91ff4d82cca40d0554f8a4d0f5978921291cd

Observation 084fcc2b-0416-453a-88f8-c2da350228e7 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.651082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.651082Z digest=sha256:565b83b9694acec1c147e446cb0b579d4438a9fb8d99fc33ec20172a53b24e3d

Observation bb3b1a12-f7d2-46a3-a3ef-2d507a968d3b · outbound

This paper cites G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.682876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.682876Z digest=sha256:87ad81efe735f1d35a829aa891282139826471967d77c5a83d67ea4b3a43cd34

Observation a6476444-98df-496a-8dec-752400d760ba · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.731186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.731186Z digest=sha256:c8415dadbdc66f50e62f728752d8bdb3bb13f3bebc86f2eef35bc079eb872b08

Observation a7c3b56b-2384-4993-af06-ca9fde4930fb · outbound

This paper cites Vision-aligned Latent Reasoning for Multi-modal Large Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.783675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.783675Z digest=sha256:5b42c2a041b5178ec07bcefe5fea5bc911717596ccbc4c6c776033d97a67a12b

Observation 420ee94e-dab1-4a97-b330-7b95dfc9d6f4 · outbound

This paper cites Cambridge University Press, 2001.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2001

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.508886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.816106Z digest=sha256:941beda2a90fe7db172f6f27e25053a21d288a34829bcfae47719e665b68fc4c

Observation a1e511ba-3ad4-4ca0-be74-f96a6923cdf0 · outbound

This paper cites Cambridge University Press, 2014.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2014

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.837244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.837244Z digest=sha256:fc1123ff13bc46424b7f21225db24d55253d91076d6720469953a12c7fc23a02

Observation 58daaab1-74dc-4075-be49-655b718d8179 · outbound

This paper cites Relational knowledge distillation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Relational knowledge distillation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.847346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.847346Z digest=sha256:ed286464c366554c031fb8ae9a5bbff5e97c3792cca60afe9e74a96069569948

Observation c5c2a94d-f3f4-41d6-824f-99fc63b9abef · outbound

This paper cites DINOv3.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models DINOv3

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.881926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.881926Z digest=sha256:729f6ca4f84e68f87f176b04fc6b20ba12bc1a4a967845710f9ef6a35ce4e4c1

Observation ceadf3e1-0473-4f8c-8652-8cbde385cbba · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Perception Encoder: The best visual embeddings are not at the output of the network

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.893432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.893432Z digest=sha256:81885cdc415b64e5c0603254ae1def6fb7ed0c1f1a992559ec5dec000dac0c60

Observation b7f2d45a-2f0f-4047-ba08-e8d01fcba513 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.442289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.897760Z digest=sha256:49b8635292b6c7decd6bc5a6ec9e127d093b4dae5ae271f34b1bce79e81627c6

Observation 7d9ab556-828c-43a9-95f8-2f452d8dee24 · outbound

This paper cites Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.383300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.927333Z digest=sha256:b1a9386b06662dbb3a8dc2b855fbd15b8c81cbfa769c6a695567419a5b84e011

Observation ae99a4e1-13d8-4c11-97b7-8dc38ff6acaa · outbound

This paper cites Moalign: Motion-centric representation alignment for video diffusion models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Moalign: Motion-centric representation alignment for video diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.272471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:28.983429Z digest=sha256:7dd3e3dc348294de2d9c0d45685bded5b3539e5cb5ad2be54f9018b1a9e07d0b

Observation 34256c33-f0dd-4f2e-83d0-80391481943f · outbound

This paper cites Lever- aging vision-language models for improving domain generalization in image classification.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Lever- aging vision-language models for improving domain generalization in image classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.153761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.047125Z digest=sha256:5fac14b580d9accd5488a7d38fd5600e066d8050835d2c8ae68e135937fcfd90

Observation b33995ff-cbda-4283-84df-d2ae3fc1b135 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.100129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.100129Z digest=sha256:1514715d992cf408f64f1d561928ade282843dbf5d28808acfdde8c2686df463

Observation be2c11c2-5218-40a3-9139-f67c19d73469 · outbound

This paper cites ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.148328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.148328Z digest=sha256:4c2c8075fb9a0ad47ca0cbff1ea1fb990ee9c43e401e004e246163688e5d7e00

Observation b9bca34a-8a20-46eb-b56e-c75c6970ab98 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.171413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.171413Z digest=sha256:8537fc09c85bbfa30d6af233d20d3877fadcc92927f5d63397ca3560f8a77829

Observation 8a819d40-0fd1-4343-a531-38dee812edd0 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.175284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.175284Z digest=sha256:394df0fe1d5eaafe9382f9f2c8816a66b02be05adc585fda77bb581d97aaa94f

Observation f08d0c3c-1a81-4e3a-8b09-b52449f00a60 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.181722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.181722Z digest=sha256:9cd6ca500bf61bf401d7b724cc7583346b20c05af1719aee46352059bf6dba61

Observation 7d3efa8d-16ba-4abc-a9bd-0e210dd59505 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Grounded 3D-LLM with Referent Tokens

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.184264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.184264Z digest=sha256:9a7523b74aa6c66329af556a5e9a667c574d04a87da21fee7d183bfeaf339599

Observation d48dd1ca-9762-44c5-a86c-d64e53f01c20 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.188680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.188680Z digest=sha256:a13b142cf6749df0680a0aa6cbbe1a9d1f878ae99614ba26dc4e0f9941994f40

Observation 80814b0b-2ab0-46c3-a529-d5a44720240d · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.215429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.215429Z digest=sha256:97bb6c953d328f7be6627aacbc3fb94e19153056ead8e58960c8b738997c1177

Observation 8f0e4626-515e-4a87-82c9-1358dc631503 · outbound

This paper cites Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.277581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.277581Z digest=sha256:7816f46714a33afec79cb51800805e7f1f6b2de428182003d01590f4688aee18

Observation 08779163-fc63-47d5-92c3-8e12b46253ba · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.320043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.320043Z digest=sha256:df82d8b8d70bcb81c82c8528848b011a6a93f27b3e69d92dff1a99256e349963

Observation 3a337eac-a6ff-4da2-9b54-bef0828e8066 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.367094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.367094Z digest=sha256:3826740433f8f82c61398751182f12d2652d426f1c6a5d3c7581dd5f9390e564

Observation a34fcf9c-3f7f-4b4b-b317-abf2803f0958 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.383190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.383190Z digest=sha256:abc88dfdd88fca728a1ff7204df51b322627cdae40d95893b16a820848173de4

Observation 42c428b8-ffda-4025-99c0-b416bcb1efcc · outbound

This paper cites Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.389249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.389249Z digest=sha256:27f0c7bdeea0fb4164f23fc3fc6f2b14d5f74bbe0cd16b7c124e32b7de6d5572

Observation 5a57a09e-1a22-4cf4-aabe-9ef8fe282024 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.392263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.392263Z digest=sha256:2e41e0ea30f461b3d8c9860865f74722cfb434ce746d9fbb56d80724b1c48682

Observation 59ee0b39-5d19-4f57-8997-2578b3ce65f0 · outbound

This paper cites Multi3DRefer: Grounding Text Description to Multiple 3D Objects.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Multi3DRefer: Grounding Text Description to Multiple 3D Objects

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.395853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.395853Z digest=sha256:d10def56b544c00bbae7acb6a02c59d422c9f40ccba431427c49ea1047f94993

Observation 8ebeea71-6527-44ba-a6a1-93d19a00a5f6 · outbound

This paper cites Scan2Cap: Context-aware Dense Captioning in RGB-D Scans.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scan2Cap: Context-aware Dense Captioning in RGB-D Scans

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.990241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.399144Z digest=sha256:68fe03a44deec1b460c7f9a6c905ac9f757d95a67bb49793a85df4e46c251bb4

Observation 8636b1ed-3519-4725-a936-864a867a45f7 · outbound

This paper cites ScanQA: 3D Question Answering for Spatial Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanQA: 3D Question Answering for Spatial Scene Understanding

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.939265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.425799Z digest=sha256:c9492351ea7a5410edf927aca2ce893a9abc7c980e77c5f5fbda9041bc89e880

Observation 32cd7efc-d531-4675-b0dc-53185f7009b0 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.481703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.481703Z digest=sha256:bb44e15f11c95023572c56bbb7c09bc12cd89e13d96e3db60c064d1a19925786

Observation 42ce05c5-32e2-4ad2-9e09-cdabc4cfc9f1 · outbound

This paper cites Mask3D: Mask Transformer for 3D Semantic Instance Segmentation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.538485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.538485Z digest=sha256:6018182043f67ae4e8184546285a57ad3390083a8dbff076d78d099b2290758b

Observation f6c8680a-2431-4bca-ac4a-f9586d1bbc40 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.122134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.598617Z digest=sha256:48f628fa2ff0bdf4c7d88b6c32a73b667c37a40d05aa469cc670b486037dcece

Observation d4b87380-ad14-4e41-9610-156e596fdc85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.075624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.662827Z digest=sha256:26015dbbc8df3d70aaf33ebc1249bd7ce7a9f66e5bf397078c9f1b85cfaf3dfe

Observation 282e49fa-6365-421a-adb5-a8656ab047ed · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.938959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.699550Z digest=sha256:7754463eb8e6c86e0f7a9b1b7cbaa35cd9ed6a9941441d507e3ee038c20670f0

Observation 16b01bc7-69f0-4b7c-8490-4910cfaabd85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.824082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.734512Z digest=sha256:c6ba41add9ee7f9ca20830e0f16ed9492301ab17f96ba4771a15d4d9a62b175c

Observation 7b4a6676-c5b5-4cf8-a6a7-226466578864 · outbound

This paper cites We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.815753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.739770Z digest=sha256:e73905518648bf55e8bdf444d2429c9558936ef6826cb243f7d26b87581abae1

Observation cc25501e-7ad7-4c4b-a7c8-1fa287ff6a73 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.807455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.742791Z digest=sha256:62fbf5e409ce3da0faf3d9a1e4303a0d339de03d8191e7767405aa60eb5f4eac

Observation 34ded46c-8425-4293-aec7-003f99db8477 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.730339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.745742Z digest=sha256:e1014ad92605af8b8018d1bfd5e472381e853aada1469484063e872a88ac51d8

Observation 79576ef0-f845-47b8-ae6a-14427c427d17 · outbound

This paper cites Dist., Room Size, Rel.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Dist., Room Size, Rel

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.611920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.749409Z digest=sha256:d35b5c5f26eff34aae66da80b67a2c2a1f7113a31319fec6fdfa5c21e1ba934d

Observation 6aed1b9d-374a-4b8b-8a21-a7fb85a65085 · outbound

This paper cites Feature. Dist.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Feature. Dist

Reference 86

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:33:30.561614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:33:29.808181Z digest=sha256:a6f70e32ec3566a9b0b076b54703e7ad84fa81c635b856690440fd2c95ae2dd0

Pith citing papers

No inbound Pith citation observations are available.