Pith. sign in

Paper Citation Record · LEDGER

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2608.10864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10864 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:33:29.808181Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54a1a48e-541c-4a26-99a9-ac63c5153957 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.118241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.118241Z digest=sha256:0820dd1e97b33aa3bf6cc970f5b071d0384db0b3bc95d1fe07ba3cc27dfab526

Observation 1f84e78f-dfcc-4e6e-8447-694e34bbad57 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.168523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.168523Z digest=sha256:d6a75d651713c8c4606448c9f167bfcb864524e988946512f36b46c7093d7529

Observation 72786601-5a8c-4972-bbd3-73db14d1f467 · outbound

This paper cites CUP Archive, 1967.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models CUP Archive, 1967

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.208043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.208043Z digest=sha256:d0c23115639afa77f5865a6b02432e51af6a60090e895d92885b54c957038cbf

Observation bb9cd8f1-1932-4808-9c04-adda660b79d2 · outbound

This paper cites Henry Holt and Co., Inc., New York, NY , USA, 1982.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Henry Holt and Co., Inc., New York, NY , USA, 1982

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:32.032179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:27.258250Z digest=sha256:22144e5d7d8d30370236ff333b6cd674ac6ade1c109e8b6a0f8708c822eede1d

Observation 8b75e6f2-0426-4b6a-847d-ff22eaf69d19 · outbound

This paper cites Number 6.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Number 6

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.296641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.296641Z digest=sha256:5ba1c39efbef37aad22d7dbf9b63f03f02e11eeca048a9acab34f9a16be0f219

Observation 9c1485a0-d1ed-4f5d-898d-49463389cea6 · outbound

This paper cites Separate visual pathways for perception and action.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Separate visual pathways for perception and action

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.334241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.334241Z digest=sha256:1556a98f68180dc25582fe166c27d01aa2121ed828b77a2861fb8bdfb205a36c

Observation c70e74d4-ee1f-4bb0-a477-e732b8540558 · outbound

This paper cites Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.374416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.374416Z digest=sha256:317d5fdf00630956b3c1bf0c3823ac0e7c562f819ec1241f66b8f916ce77c184

Observation 53fa584c-0db6-46e2-a28b-fd03061b08d9 · outbound

This paper cites Oxford university press, 1978.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Oxford university press, 1978

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.403655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.403655Z digest=sha256:5ca3a4a9355205ea6d630b6fbbc53ee001a819ecf9f51a9162a08ea1dc7b94b0

Observation 3d47e3f0-36cc-4b90-b178-965802fae45f · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.434416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.434416Z digest=sha256:c68e1a6f3f421879df07b1613614c1312896f5ca8fbe01a6ec59f347b2919311

Observation 04ee055b-8d94-4616-9a27-fa3eac91f6a3 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.467684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.467684Z digest=sha256:ccb717849b3960f9167a5833496d7ed4e36f283af5bdd30b7b98a3d949551a74

Observation c274fcc8-ac9b-4312-b3d0-4d118e2a2623 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.488894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.488894Z digest=sha256:2f54eca72ad28d950616294ef877fdaa9f2f481f96c4a6752d1540d92e78d8d0

Observation f4020c15-82c7-4b85-b0ef-3ff96bf5ab12 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.513243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.513243Z digest=sha256:8dc4af5be3c9668d7cdace888ef8f3dff52ca551dfc6ea5c2155febe004adaf9

Observation ecd75291-ac67-40ae-9e16-39afc05ac79f · outbound

This paper cites Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.536117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.536117Z digest=sha256:f8ba725b6dfd4ed86d46ea76866cf89d5a22cdaeb0bdfa4d5c6395eb04e76be8

Observation 0f5fb8ed-eb7f-4d78-a561-d48b5f067980 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.559154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.559154Z digest=sha256:923291e630277ff40058def371fc81fffc4bea6598165206c19e3004e04aa335

Observation 205a6793-af32-450d-9f26-0bdfd8316df2 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.597794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.597794Z digest=sha256:183603d14359ec4f655c26f7dc9b25ae45c750fe5550af8989a4c9f15c407c2e

Observation 61b7547e-1d9b-4a8b-9164-36bb4225850c · outbound

This paper cites LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.861793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:27.627853Z digest=sha256:4a0c3096f5a9df08d3e8e290cf14a312b86be42574605dfee262fc91ed98481b

Observation 7bb53d38-73a0-4035-a903-7194cffba8ed · outbound

This paper cites Probing the 3d awareness of visual foundation models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Probing the 3d awareness of visual foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.851699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:27.662301Z digest=sha256:a0bdd55114cd8568cc140c39048e0754a536030453f3bc377ce10ab38e45ed51

Observation 5cf0e01c-42fd-495e-a90d-a30d3d2c7c98 · outbound

This paper cites Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.693368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.693368Z digest=sha256:2bcc30be18fc6e3a43e5dc1106ab8f5e8b58b202b608e77948d5be41b8bfc1b8

Observation 87d6ce1f-049d-4151-902c-6f4b92e319bd · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.750364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.750364Z digest=sha256:a767f21c917fe7ec3e9e2d63c7d1b835771f4fa55b1ba02784da599cb06f7253

Observation ccddb12c-aba2-4058-990e-f3875f62cebe · outbound

This paper cites Vlm4d: Towards spatiotem- poral awareness in vision language models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vlm4d: Towards spatiotem- poral awareness in vision language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.802712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.802712Z digest=sha256:882d84a4df7d968de5b70f36715bac198b3a53782a4065382acc0e0142bbde48

Observation 4a06d401-c430-4d9c-bbd9-3f0656262306 · outbound

This paper cites Cambrian-s: Towards spatial supersensing in video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-s: Towards spatial supersensing in video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.915456Z digest=sha256:88219250a65d955218218a73815c7d681ad29f3486a3cac141b292225b51298b

Observation beac19f4-b222-40b6-9e7e-69cf97e0fcbe · outbound

This paper cites Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.973683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.973683Z digest=sha256:a0d90433c62a52bbf0241ca3038ff0116f0cfe1ec2dc73d96f133eb44238ce3b

Observation 4647a97b-0e1f-4c31-bcbd-a85f98f5f94e · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.004225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.004225Z digest=sha256:ebf207c79d328ab79bb589835077c4251ae5e370d8ebb0e493d96ebbfff1f6ea

Observation f434eff9-3b11-4021-8934-ecd778e162be · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Distilling the Knowledge in a Neural Network

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.057213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.057213Z digest=sha256:6902e42e3010df3080b48d0e709e562b7cd8b469f58ec6e8126df8940ba6a139

Observation c976a4dc-f4a9-40f3-b939-dd1c9a60943b · outbound

This paper cites Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.803965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.107936Z digest=sha256:bcadf0cc29bbf17b87230ab18e1adcfa5ad1c73008e15b9a0264047deb2f8607

Observation a77fa326-3b6b-490e-bc6b-4cf2003d5c93 · outbound

This paper cites Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.688307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.143957Z digest=sha256:72d16d5e928997b65d51cc09f0dedd1219a8384f3ec2ea2f522c90d5432a3685

Observation f1a1f640-c5ea-4332-b960-e17a12b4f2f1 · outbound

This paper cites The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.182390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.182390Z digest=sha256:6496a323dfa9285a13238c2489ca0a69570b4321dc03cfa6e4fa2d1300f2d46c

Observation 0b40abdd-a5db-4460-8a8c-2ef393025aae · outbound

This paper cites Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.212915Z digest=sha256:303f25742d4793e326bab13b837b10f4d577794d6afa1ac0101a631334bcc234

Observation a6faf50b-53c7-4948-b050-b6ca50e69b8f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.247419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.247419Z digest=sha256:00b799decb7464e6a58f8d87322c98aed0fe5a3eed442dc0b951721db792e414

Observation 67da9146-8180-461b-920c-a09a20989136 · outbound

This paper cites GPT-4o System Card.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models GPT-4o System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.289469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.289469Z digest=sha256:1e79599d33b11b902452d866722c05c4efc531989f41e07565f75e0503cc040c

Observation 7c938b5e-59a1-4a40-bc81-2a4b0f23cbd7 · outbound

This paper cites STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.317477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.317477Z digest=sha256:14486b6faf34cbf68e047dd8c5a2a968f0f498a7c21405d2505cbae9524758e0

Observation d9ce87d5-a09d-4627-9a8c-8812128504eb · outbound

This paper cites VLM4D: Towards Spatiotemporal Awareness in Vision Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.342219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.342219Z digest=sha256:425bb5d50acc93a961d9399a23867083cf672044f8b869d9e3ce5b9fb9ca5493

Observation 6ccf82cd-9f78-4f08-9861-33308bddca74 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.364382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.364382Z digest=sha256:8b06bbb3708597c7fac7f2834a6ca551ab096d6182cfb29855e63a50c0fc812e

Observation 1c083e24-1267-4b2f-94ec-a3fe365ffef1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sigmoid Loss for Language Image Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.381618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.381618Z digest=sha256:d9f88c92f6b3044a9394825e81399b0a87b4e7aaa60cec2b0e5da2e1de7c3ec9

Observation 8e3a871a-0a6a-473f-a70b-2be1a1d4fa0f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.399065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.399065Z digest=sha256:e1234a1da8d941c851472bb843d0aa6096dcee0ac8cf78b1515d38c9ed191940

Observation 96ab4e6c-53d2-426a-abc4-f0e554584a1f · outbound

This paper cites O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.460314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.460314Z digest=sha256:c3647d07033867283143c71b9bea1a50adb628f1ac90a8d95cdd5d2337b918cd

Observation f65e9dfb-4032-429b-9eef-45796e4626c8 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.508799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.508799Z digest=sha256:0122f7461459c6d86c3ed08112c384a860d7a2847126ef395b9e7403edfbd4f9

Observation ebcc31ef-8308-4671-a011-7d8922eaf6e2 · outbound

This paper cites Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.562622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.562622Z digest=sha256:1163b0d17ede824d800d7ccbd24740cabb4d986241ccb6713c75d8b858455fd9

Observation 06b1887d-f586-4ca0-8901-29e9cdbf6dd2 · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-S: Towards Spatial Supersensing in Video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.585047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.585047Z digest=sha256:c4fbeacf0c38cf79d7a6eb5328d2c88b7705a2c9ecf4bf3f939aafcbbfd53dc9

Observation da5f7c1d-cf31-4b45-84aa-571d67dcbbf6 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vggt: Visual geometry grounded transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.614240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.614240Z digest=sha256:dc5551dc39a81793930ea4bfb6b63a5b9afaab97f3e35affc62837d3da7beed8

Observation 10bccea0-33e6-4810-9190-2e9cd6e583d6 · outbound

This paper cites Continuous 3d perception model with persistent state.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Continuous 3d perception model with persistent state

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.541921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.641131Z digest=sha256:dacd01556c43d232c73cedb37466d2cdcd44b0a93fd75aa3e2f72b8daf539461

Observation 084fcc2b-0416-453a-88f8-c2da350228e7 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.651082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.651082Z digest=sha256:255f0aba77a61cd9a8c1847f40ea6d2ac22c8b7b5de0456009baa9f408dbaa11

Observation bb3b1a12-f7d2-46a3-a3ef-2d507a968d3b · outbound

This paper cites G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.682876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.682876Z digest=sha256:42e7722fda3cb2e40e05c5e09ffa283c1818395d8ca891690b6f976906ee6a1b

Observation a6476444-98df-496a-8dec-752400d760ba · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.731186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.731186Z digest=sha256:5ce31b50dc711483f9f7318d828e9cf46e8c1a497a07b72a03a978decfce7f64

Observation a7c3b56b-2384-4993-af06-ca9fde4930fb · outbound

This paper cites Vision-aligned Latent Reasoning for Multi-modal Large Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.783675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.783675Z digest=sha256:9605c9583bc89948d0bd670626d46e47463fb1f8df94ca5afc544328ff617747

Observation 420ee94e-dab1-4a97-b330-7b95dfc9d6f4 · outbound

This paper cites Cambridge University Press, 2001.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2001

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.508886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.816106Z digest=sha256:7dd0d8c0bba9efd5a95851125b2ff708277e660ecb2c3d45b74d17a65177c5b8

Observation a1e511ba-3ad4-4ca0-be74-f96a6923cdf0 · outbound

This paper cites Cambridge University Press, 2014.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2014

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.837244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.837244Z digest=sha256:38c3a8f3faa16bd432c4fcf0c70edb40d1b2f822c4a33f6f47593a6d9fff5c3c

Observation 58daaab1-74dc-4075-be49-655b718d8179 · outbound

This paper cites Relational knowledge distillation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Relational knowledge distillation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.847346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.847346Z digest=sha256:d7165c459e8d697c39ee6bdf9872eaa63ded468f4578b847502402252d49bbda

Observation c5c2a94d-f3f4-41d6-824f-99fc63b9abef · outbound

This paper cites DINOv3.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models DINOv3

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.881926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.881926Z digest=sha256:1139924ad8f210190d7b60f1d5002895ae0c9d6c62de83d1b0f2fc6722952366

Observation ceadf3e1-0473-4f8c-8652-8cbde385cbba · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Perception Encoder: The best visual embeddings are not at the output of the network

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.893432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.893432Z digest=sha256:4cfbba26231bbd12fb26ea808df41a1ba265ec67ddc2c7d1c00fd521ba35dc4a

Observation b7f2d45a-2f0f-4047-ba08-e8d01fcba513 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.442289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.897760Z digest=sha256:a6979c57ca961ef750cc8f082e311ad70360c2d27a909076c28f1ac5ba3d31c7

Observation 7d9ab556-828c-43a9-95f8-2f452d8dee24 · outbound

This paper cites Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.383300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.927333Z digest=sha256:1ca2e79a41aa8b8fa0345c51111acf3829689aee97ac8b1e0e819ebdbb6f084b

Observation ae99a4e1-13d8-4c11-97b7-8dc38ff6acaa · outbound

This paper cites Moalign: Motion-centric representation alignment for video diffusion models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Moalign: Motion-centric representation alignment for video diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.272471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.983429Z digest=sha256:62547670d9361a794868037d7ba1398a6214510fdb0f54e19e2fc365d7893809

Observation 34256c33-f0dd-4f2e-83d0-80391481943f · outbound

This paper cites Lever- aging vision-language models for improving domain generalization in image classification.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Lever- aging vision-language models for improving domain generalization in image classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.153761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.047125Z digest=sha256:3d47ed4ef0295e219415dd72f938084dc7e8fe396cfcc83b32d5c667ff0b2dfb

Observation b33995ff-cbda-4283-84df-d2ae3fc1b135 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.100129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.100129Z digest=sha256:d3fdb8d7f8ac2c4a851f7d51ef6339d26632b01522904a1a1e5e760a23e99962

Observation be2c11c2-5218-40a3-9139-f67c19d73469 · outbound

This paper cites ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.148328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.148328Z digest=sha256:7d1bd404ce88eaf29e1a16e8c6e9095891e9f32fa0051f9dc7acd4bc1f05f72b

Observation b9bca34a-8a20-46eb-b56e-c75c6970ab98 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.171413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.171413Z digest=sha256:847d8555b264039b4a9f99e2541c4818a5e99df43bac592b33ce32c2d47d9286

Observation 8a819d40-0fd1-4343-a531-38dee812edd0 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.175284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.175284Z digest=sha256:d67ea09323f5a7cf3625844e8a73819aee29775208ee3198d08ed8e45ab1a94d

Observation f08d0c3c-1a81-4e3a-8b09-b52449f00a60 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.181722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.181722Z digest=sha256:149ccf831e62076cb84346898d50863657b6d900c6a97d9690a1d18963e60741

Observation 7d3efa8d-16ba-4abc-a9bd-0e210dd59505 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Grounded 3D-LLM with Referent Tokens

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.184264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.184264Z digest=sha256:7ee9362a96d88f6c4343d7bdc41ae0b1f98afda5fa7324c602edec014af0559b

Observation d48dd1ca-9762-44c5-a86c-d64e53f01c20 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.188680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.188680Z digest=sha256:bed17ddaadbcdc379f66bf01faa99a82f179391dbc1c819c2af071d0d79274d5

Observation 80814b0b-2ab0-46c3-a529-d5a44720240d · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.215429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.215429Z digest=sha256:8e9191d22e37b3a445cecf6252a449f5ffad9219c444e4d93d1dfe572a12616e

Observation 8f0e4626-515e-4a87-82c9-1358dc631503 · outbound

This paper cites Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.277581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.277581Z digest=sha256:fcc546b3949e40aa700a7e1f1533caa7e4961545f470ad9f3bd65bceacc86125

Observation 08779163-fc63-47d5-92c3-8e12b46253ba · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.320043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.320043Z digest=sha256:c4589b813c8c293b3c41f63d80355dddc7a0ae0e65ddea5b81bc7ec5bc25f71d

Observation 3a337eac-a6ff-4da2-9b54-bef0828e8066 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.367094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.367094Z digest=sha256:e75927596f17a0ab4a94212d82fcc78803a071146cf0aea14a3a45a089de6de5

Observation a34fcf9c-3f7f-4b4b-b317-abf2803f0958 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.383190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.383190Z digest=sha256:1d124d91c43790aa3795791d9c527b756568a892d5f2822f866b83d8f0b7891d

Observation 42c428b8-ffda-4025-99c0-b416bcb1efcc · outbound

This paper cites Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.389249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.389249Z digest=sha256:2bd5a2073cf88c55131d2ba3e18863617e80ee4b81c9ee6755bef58e3ddb2acc

Observation 5a57a09e-1a22-4cf4-aabe-9ef8fe282024 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.392263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.392263Z digest=sha256:888a2d44b2dfeb98f541f0abf297df2c14d127a2b6f21b1bb8d4fd0914154bae

Observation 59ee0b39-5d19-4f57-8997-2578b3ce65f0 · outbound

This paper cites Multi3DRefer: Grounding Text Description to Multiple 3D Objects.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Multi3DRefer: Grounding Text Description to Multiple 3D Objects

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.395853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.395853Z digest=sha256:0423555ee22a6fb9d489ab34e36f62fed8e40123cc9600b46d3c308ec881e250

Observation 8ebeea71-6527-44ba-a6a1-93d19a00a5f6 · outbound

This paper cites Scan2Cap: Context-aware Dense Captioning in RGB-D Scans.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scan2Cap: Context-aware Dense Captioning in RGB-D Scans

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.990241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.399144Z digest=sha256:24cc4b15cffb0ffaced65fa46b97b46d1bee08651a95fcd1bca7e66737731a99

Observation 8636b1ed-3519-4725-a936-864a867a45f7 · outbound

This paper cites ScanQA: 3D Question Answering for Spatial Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanQA: 3D Question Answering for Spatial Scene Understanding

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.939265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.425799Z digest=sha256:9cb8112e36fe53324446668f8bed0e6e6ed952ef9331de9a6de7c10abdf9d496

Observation 32cd7efc-d531-4675-b0dc-53185f7009b0 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.481703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.481703Z digest=sha256:68404668ac1a1657eadad71b47802c82f9ef6b5d8a6c6feffd8ccedfcbe8dd98

Observation 42ce05c5-32e2-4ad2-9e09-cdabc4cfc9f1 · outbound

This paper cites Mask3D: Mask Transformer for 3D Semantic Instance Segmentation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.538485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.538485Z digest=sha256:a052840d20527fb0e5bce2e6eb533a228a4b617478f9577dc5dcb7e4280ab948

Observation f6c8680a-2431-4bca-ac4a-f9586d1bbc40 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.122134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.598617Z digest=sha256:d7e2e95e512d0d2c24150f06e1953c35c64cf7f09719ea073dea89ce65776b59

Observation d4b87380-ad14-4e41-9610-156e596fdc85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.075624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.662827Z digest=sha256:75608bd36ae293041ff9c387e24517a107ccac46642283bebd90afba6331505c

Observation 282e49fa-6365-421a-adb5-a8656ab047ed · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.938959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.699550Z digest=sha256:abfcb751ae2b54e8343a9207e0eb76ca90ce8a617846d745702765acebd3890f

Observation 16b01bc7-69f0-4b7c-8490-4910cfaabd85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.824082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.734512Z digest=sha256:df73ef75f3576ccbedba39313d72551a3f2ad8cd912c290112c43d84b1db8e16

Observation 7b4a6676-c5b5-4cf8-a6a7-226466578864 · outbound

This paper cites We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.815753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.739770Z digest=sha256:d8448f425461ff3477851271f301ec5da7a03d54361fe659276a74b4e341be5c

Observation cc25501e-7ad7-4c4b-a7c8-1fa287ff6a73 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.807455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.742791Z digest=sha256:20a0d4d9b101c66108ff2b696913e30f69c374ea3c0bbe7589e3c8c23316d0b1

Observation 34ded46c-8425-4293-aec7-003f99db8477 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.730339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.745742Z digest=sha256:880c9c28c441268ca44f1cdae6bd9e2dc1dcc5559cd796d5af3017c2470619ae

Observation 79576ef0-f845-47b8-ae6a-14427c427d17 · outbound

This paper cites Dist., Room Size, Rel.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Dist., Room Size, Rel

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.611920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.749409Z digest=sha256:e56810efa04726fbd719a7b0688b406a8ffb00a18103648efa5d2e828906f017

Observation 6aed1b9d-374a-4b8b-8a21-a7fb85a65085 · outbound

This paper cites Feature. Dist.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Feature. Dist

Reference 86

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:33:30.561614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.808181Z digest=sha256:592123e9c1214e0d1de15fca91fcb6929f4da4fca87ba61026420724a6bd33fb

Pith citing papers

No inbound Pith citation observations are available.