Pith. sign in

Paper Citation Record · LEDGER

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

As of 19 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2505.00788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00788 v3

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:40:27.818667Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:42:12.968388Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 406a59fe-1e26-4614-8284-1220506dbdf7 · outbound

This paper cites GPT-4 Technical Report.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.528694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.528694Z digest=sha256:12f8b6399689608471fa7274af0f5837162c99681bd77c43f57f8935d4195b26

Observation 3899654e-51cf-4047-8147-8fc65620184d · outbound

This paper cites SpaceLLaV A.https://huggingface.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpaceLLaV A.https://huggingface

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.715270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.533093Z digest=sha256:2c0de25effde5d3ed87fcb6084b7de2530404fe1a8cdbc3075e9aa91d6114a43

Observation ef159a4d-37d4-4506-bfd3-d6e4a9577b20 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.537060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.537060Z digest=sha256:0796ea8efb3ce3c210ebdc5bffa15ed0813c66628fff46666fa39f4856806605

Observation 54101c4c-4803-4e16-8d9f-cd5527539fa6 · outbound

This paper cites Claude 3.5 Sonnet.https : / / www.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Claude 3.5 Sonnet.https : / / www

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.691472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.540724Z digest=sha256:9399be221614fe2ebd1395bf8ebac73e477b3254cc1e03fe8df1d1f27d037963

Observation d08d78bf-58cd-48f3-a510-6f425e17e097 · outbound

This paper cites Apollo syntheic dataset, 2019.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Apollo syntheic dataset, 2019

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.676342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.544404Z digest=sha256:88ca4fccc83d9dc29ba486df3c7b753c6727dcbfb907395b06f758549269d18d

Observation 1c09667c-7bfd-433e-beed-de79b75fe683 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.662827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.548786Z digest=sha256:5ff1fa3dd744169d869606aef479baadb7367f1b3037f7e049c2b6104aae43a6

Observation b680ca9b-60bc-408e-b398-2ded5eba2dc2 · outbound

This paper cites Qwen Technical Report.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.553190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.553190Z digest=sha256:82b513834a1d678da278368c1330435d90b57a1f1182b2c9557561632a2c822a

Observation 210bf3bf-22e8-47c5-928b-c18b4892d11b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.557935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.557935Z digest=sha256:016226496a5537620ef45bf4802e3d16b4a69513c4fe986164eae497e78f98a5

Observation 014d3939-6660-4473-93c9-d6dac68e29ea · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.563585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.563585Z digest=sha256:911f1f321a5a0cf0a75b321b7a71d88b73b76f90267c701a393c0c54577c032f

Observation 2379b6d3-b0e9-4702-8586-5485bee21225 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.568854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.568854Z digest=sha256:2b0741f5cd8fd8d37097dc4c4766f869cdbd814e8e3241753763bfd812b9e40e

Observation bdcc0a7c-ab19-42be-8fca-2fcc0b0824dc · outbound

This paper cites Omni3D: A large benchmark and model for 3D object detection in the wild.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Omni3D: A large benchmark and model for 3D object detection in the wild

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.649771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.573588Z digest=sha256:223f1c65ec988319285dd56a64e225bf7b0f69efe44a367db9bfab6b88007aca

Observation f7518ae4-2fea-4820-9626-77596264aef2 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models nuscenes: A multi- modal dataset for autonomous driving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.634820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.577671Z digest=sha256:87336062a9f4130282432a2e08f9567f6bc7e3873420634601447a495431a56c

Observation b3b2ee18-f9ae-4442-b275-15616a4374b7 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Honeybee: Locality-enhanced projector for multimodal llm

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.621016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.581845Z digest=sha256:7e3d8317097bb2a8dc1a85eef05f722a55d4b3540725a2d83bedd16306e6ced5

Observation 6613005b-fd73-4555-b321-9c286ba2a0c4 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.586160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.586160Z digest=sha256:053ffedeb65fc382ba4b9c5e90fc6478f24529d13d597510c4fb064744452bcc

Observation bb75f2f4-66a5-4a0d-b871-a3b9d668f78a · outbound

This paper cites Vitamin: Designing scalable vision models in the vision-language era.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Vitamin: Designing scalable vision models in the vision-language era

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.598273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.590245Z digest=sha256:41e1ef467b3787aa183b5190b8b02e56e79bec5ac50db0ba5224a6ad87ae559e

Observation ee237961-8d51-49cb-87d3-00b205a249fb · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.594324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.594324Z digest=sha256:d4d708dd362b3f71f961c2339f302be9b26eab3509865f2a34790f62fc385a48

Observation c7428130-25d9-401f-b0f8-ff71c63dfd4e · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Imagenet: A large-scale hierarchical image database

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.584277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.598798Z digest=sha256:c558c90348bda63b801013e5ace1f9a61e11fb86d9e816311867389198f1b89f

Observation e2fe4363-1272-494e-bdbc-866bceb42454 · outbound

This paper cites The Llama 3 Herd of Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.602876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.602876Z digest=sha256:e525b3fdf79d4d793e69c14e1611f39f234940b60e67e3fd558fd3322790562a

Observation 8f4a9e3c-3739-4cf3-b85d-a0bcc6ded7c3 · outbound

This paper cites Prob- ing the 3d awareness of visual foundation models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Prob- ing the 3d awareness of visual foundation models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.571158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.607343Z digest=sha256:2b576721833946a91a9b84336ffe30ae00cac3a61d0e2477f11c7731d93a8e76

Observation 01858a9a-0c1a-4c10-9b5b-b5316d65ceee · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DataComp: In search of the next generation of multimodal datasets

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.611469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.611469Z digest=sha256:da0b12eb9acc6f3914629220b4faf4caa1364e20717e9b9099fac437d82ce988

Observation bf252cec-806e-40cb-bddd-14c085680d2c · outbound

This paper cites Are we ready for autonomous driving? the kitti vision benchmark suite.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Are we ready for autonomous driving? the kitti vision benchmark suite

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.558228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.615737Z digest=sha256:6ab02c5d3f0f0ccb3736bdfe6b3595817102a3f5e1c9f4487416b0e059b54c2a

Observation 963710cf-9f38-43ac-9697-a4faa7d1eb9a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.619809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.619809Z digest=sha256:2c2173cc3fd8a4211b8ffeac4971844119d68ac8437d51a6bd57af31545bde5d

Observation 5853c801-4a09-4244-a610-be3722600d4d · outbound

This paper cites Masked autoencoders are scalable vision learners.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.535650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.624025Z digest=sha256:67b51ef2c3556394cd763de8f42af54268578df24e38a4794bd3f38d20ff7856

Observation 0bb241a8-9b5a-49e6-a74c-b4964909fd04 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.628147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.628147Z digest=sha256:9b523bc550c8589b4570b3554fdef633121aa0e1f2be7a558fd3e35fb7d4e0b8

Observation ab3e944d-9ffc-41be-8881-0afa42753c6a · outbound

This paper cites Open- clip, 2021.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Open- clip, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.632345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.632345Z digest=sha256:c2914ce044ebb310f45564f5b6978fbf3397bb1de5e3419233857c01dea619a5

Observation efad3a13-2355-4f59-857a-bcf21ed9e00d · outbound

This paper cites Novum: Neural object volumes for robust object classification.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Novum: Neural object volumes for robust object classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.636714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.636714Z digest=sha256:d31df88e2f8154faf85dbe321da2acb96cd013543e45e9b3fa4f121dfccb4e5c

Observation 0a18a1e3-b8c5-4f12-919f-d9372cfeb9b4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.641122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.641122Z digest=sha256:9519f8044662d067a88758a1a31dd1897d6780b3a9ee2c31f7d2eb1326e92b2d

Observation 8b70a587-e6b4-442f-a52b-039d3219ebb4 · outbound

This paper cites Mistral 7B.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.645314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.645314Z digest=sha256:c5db5d9e0a0ed41ae04602ba5626e895e496ebffb5131163d54491d400d181f7

Observation 23def1e4-6e16-456d-a272-19374143cbe0 · outbound

This paper cites Perspective fields for single image cam- era calibration.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Perspective fields for single image cam- era calibration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.496761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.650512Z digest=sha256:658d47bd4472ee90e905bde5112912eb7acbc9f75cdf87cbf2641b053edb2b27

Observation 0f29b4b2-cd60-4d8a-bccf-fe8980c20c1c · outbound

This paper cites Segment Anything.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Segment Anything

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.654030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.654030Z digest=sha256:2b4f57e587cc09da61a1ea2c63f2c17176e6b24a8fbdf4df42d6fec62c8f1134

Observation d6ad7280-ab4e-4084-8836-82d3c4087e82 · outbound

This paper cites Segment any- thing.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Segment any- thing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.482808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.657739Z digest=sha256:67be436984815cc155c0ca8718a7bbb8ad0eeadb6e530e292f56139ae9d0333f

Observation e8503dba-ab6d-41a7-ba90-27a7802de402 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.468752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.661340Z digest=sha256:852c6ee9e4d66d82883cea453d4f97f17c3c270f0c78bafa9b1af79bf59cf421

Observation 6133f4d5-e8dc-43b8-b982-5d7bc1ba9234 · outbound

This paper cites an unresolved cited work.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:40:28.455596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.664986Z digest=sha256:ad8078bec7a586edc552e1157a364304b4b5c23660e12c0a08687b90f4c5956f

Observation ec3d869b-35a5-4bf8-a173-9680cbc1cfd0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.668620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.668620Z digest=sha256:e9d1e79b30c5c148996d0e61dbdf41e849e1a7b3bd17d522b2b1ee2c0a57b2ee

Observation da067013-52b3-4923-b835-b331d5dff1cf · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.672593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.672593Z digest=sha256:813bbb1be6de73b7a9f1b89fcfc7f3092849b6c70b96d6433a73e8e134196965

Observation 6a9eeaed-3940-4f5e-bf21-6cac2194feff · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.676116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.676116Z digest=sha256:5ab1a443baf953b1ec03faaa5597f45625383b8e02995516673183cc4cc460ea

Observation e191db48-ad61-4223-9dab-3e802c457a8f · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.679528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.679528Z digest=sha256:8276ca0614cfd04597b5dc799a511fcb07a902481d26bbad1d932ca279d4fd19

Observation 2d8537d9-59b3-4597-bf88-eb1310a983ab · outbound

This paper cites Learning customized visual models with retrieval-augmented knowledge.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Learning customized visual models with retrieval-augmented knowledge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.424103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.683992Z digest=sha256:b8dd616610758f27f303310a39c75b0cce15ce194548b8ecd7ca3e7811cc8a7c

Observation b5370f09-d4d9-4e65-823c-dd477f4ebdd9 · outbound

This paper cites Improved baselines with visual instruction tuning.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Improved baselines with visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.411147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.688146Z digest=sha256:3a5e1a0aa863a70ec01cef8279a757cc5d663f59a6894937767e0b1f913cafe2

Observation 8563a318-bdce-4f59-94e2-8c66228b8f8b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.692491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.692491Z digest=sha256:386b7a55fd902ff973eb418b4519f744643d519cc4065fbc20a9abab3bf35dbb

Observation 9078385e-9047-431d-88a5-8539b76bdac5 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.389689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.696559Z digest=sha256:7d925ec300869ebd0af47d646b26ae01c63469401f053077d6243ba13475d2d0

Observation 2d40ee3a-6590-4210-971b-7cb5e430e012 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.701110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.701110Z digest=sha256:110430c84e5470705c2f0cf769d5b80406c427de041e68a05c75f7062a5a9efe

Observation 1b4f0ab1-d640-46cc-b582-d973fb3af87f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.706213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.706213Z digest=sha256:247cd75b168c2f68597d5f1f0682742c08dce33a5375a8ef9830bfb81bb36f45

Observation 9d41879e-6ea7-45c5-9dd6-5fb92975967a · outbound

This paper cites Robust category-level 6d pose estimation with coarse-to-fine rendering of neural features.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Robust category-level 6d pose estimation with coarse-to-fine rendering of neural features

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.710906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.710906Z digest=sha256:26b962882bdeaec4bacaf7ced6521267cfa6a6677128cb74fedeb920909f8410

Observation cb3e3658-1c82-496c-8a94-595420f171ad · outbound

This paper cites Imagenet3d: Towards general-purpose object-level 3d understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Imagenet3d: Towards general-purpose object-level 3d understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.367631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.715210Z digest=sha256:5bc6d3bec5e6fe42534353b1492df71dc93b3bfafa2b55734070ce4a873ae99c

Observation ee39a5d8-b409-4f20-afb8-66877bafd94d · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.719431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.719431Z digest=sha256:9a36b0a49ca9694196e21ecba9aaa0e709d4066459058d5b33b75ea3199897e6

Observation a0f212ba-c183-4ba0-ac22-e03927a53aff · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Dinov2: Learning robust visual features without supervision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.353857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.723958Z digest=sha256:2fa1603d2649289fac65565698b2281b22b829bb8e306d43525a2a92953df032

Observation 46c97bf3-ed26-462e-8532-8f29b56a2c81 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Learn- ing transferable visual models from natural language super- vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.340434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.728221Z digest=sha256:d2e5003de9a9dabed06b6a363939f1673c216f3e82154a03c22611f8ed97e1fc

Observation ab98a02a-21ff-4a25-aa69-358024560e93 · outbound

This paper cites GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.732388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.732388Z digest=sha256:becf162ba7e6969652336512061ac4d434cd403ae02bf20bf8b54d011bc64b50

Observation 5ce633a1-7150-4fe4-b5d2-0c05685ba58f · outbound

This paper cites Susskind.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Susskind

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.737271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.737271Z digest=sha256:92efd716e11e1013a13240b37883377858a06698c36fde21a7058b96a8adbda3

Observation 0480b152-9d4f-456f-8a1f-4cf1911f0bd2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.742087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.742087Z digest=sha256:18723a6c2f055011b6648c8fce846bdef9e1b842bba8f17b1ebdfc9f513c13dd

Observation 006893b9-ae58-480a-858a-34eae99b0e66 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294, 2022.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.746271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.746271Z digest=sha256:bffad7bfeb2be9505fc046bdf92f4686ec4560fe1dd8b679711fb80829db2754

Observation d9eed4e5-c451-414b-9f1c-d369abc0d3cf · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.750251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.750251Z digest=sha256:5e397e141e7a4a67953704d709d9430e17c0251ef3f00ffbb6b896280efc2e3c

Observation 38d7f83c-8884-48a2-a43a-5028083c44e9 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.292913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.754431Z digest=sha256:da5140e8d04c5329cd67e52aed891a4625588e9b713142709f73b469a00f1e64

Observation 5079e0ca-85ec-40d6-8dd8-fc7220e45669 · outbound

This paper cites Core knowl- edge.Developmental science, 10(1):89–96, 2007.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Core knowl- edge.Developmental science, 10(1):89–96, 2007

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.279345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.758501Z digest=sha256:38e13fb063a0d27889209703190a426c530e17094204d2aa7e219a1ae0817eea

Observation 2d34e57a-31d7-4600-b65c-438238908d63 · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Revisiting unreasonable effectiveness of data in deep learning era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.266026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.762558Z digest=sha256:ba3ae86accd9ae4940e2800ecf8609b6f38c2727e820c58207f3002da12bf4e1

Observation 136a0093-736c-4ee3-81e0-c392bb18cacf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.766525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.766525Z digest=sha256:8c37fe8c269a0b9520cd986aeac2b835a3b61a6a4b620b61b8203398da23aa5c

Observation 7432001d-b856-4980-8200-07a6a7a57b7b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.770707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.770707Z digest=sha256:933f2574bd16284b4ae2148738f620064f69cef0b42b49bb8042a1767e1410f5

Observation 58112f68-ab1b-4fd0-9de8-9a674ab0493b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.775308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.775308Z digest=sha256:97eb7dfc0bc61b086d0c3d78444f399e2e28405706628e730cd3398572db3df5

Observation 1cb36f74-af69-48b2-b64b-007267717ddc · outbound

This paper cites 3d-aware visual question answering about parts, poses and occlusions.Advances in Neural Information Processing Systems, 36, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models 3d-aware visual question answering about parts, poses and occlusions.Advances in Neural Information Processing Systems, 36, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.252560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.779103Z digest=sha256:0c5d43db1758ac02c36a86bd2b6bbc272232b3bd7a8b324584bb2b1d46b1c152

Observation 87ad28bf-6065-4e22-a3aa-3ad93f02c52b · outbound

This paper cites Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.782481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.782481Z digest=sha256:fc29022ae6328bda79d99d646d2b39f19a7856cc6057bfcc68df8a49203cba5f

Observation 9413468e-6207-41cb-a66c-0ba8a4f1ed2b · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.786283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.786283Z digest=sha256:f2752b9b51a015d3918421c46eabebc28ba8ee947dfd62212011f09f2b431fea

Observation 5bcdd412-64c8-4b6c-bae0-b479654111c3 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.789631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.789631Z digest=sha256:99ab9b387fa6f247db40a6b8f1b707a12d95fb7e9b2b98c446e70740df111438

Observation 641aa501-4832-4c54-a095-8c145200b324 · outbound

This paper cites 3D Question Answering.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models 3D Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.793012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.793012Z digest=sha256:5c2e6de0687c3c733bdae52a3c8a69c25131cbfafc0397b7e47166aae31ebee0

Observation 0a28dadf-2403-4a41-b966-8c54e4654bec · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.796715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.796715Z digest=sha256:005e807d2f52aae58e8befa6c3dae39042b849ad5bedec3ddf02aa54951d8034

Observation c0976f13-42b8-4a8e-b0c9-79e2abede03f · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Recognize Anything: A Strong Image Tagging Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.800978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.800978Z digest=sha256:4899def098b6160c14069afc06f0e991375016acea3724d38de7d849d341a572

Observation 05828b61-7e53-4419-969c-22fdbc6382e6 · outbound

This paper cites Recognize anything: A strong image tagging model.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Recognize anything: A strong image tagging model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.222985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.805680Z digest=sha256:42a0a8c944dde33548e3b5ebc2b9031f856024e52c515bb7494beccffce76b01

Observation c6ea85e2-6324-4c84-89cb-3e5cc5a09a3b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.210551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.809925Z digest=sha256:cfc725d5c6b05be22ac9d7d1e717134d785ac641716662d3a776926a6c9f17f2

Observation 3a2932da-f0a4-413f-a27c-6f50635ac041 · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.814069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.814069Z digest=sha256:745ad12eda77b8c793e726af3cd1204036635f96b7c9ec25af4dcf4f3b189679

Observation 40f19090-a735-4d0d-a78f-2b57a9580da8 · outbound

This paper cites yes” as the answer and 120 questions have “no.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models yes” as the answer and 120 questions have “no

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.195914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:40:27.818667Z digest=sha256:eeb98f789ab9c2e86f1525a3c21b879df5ae1508ecec2e3b9a3c2145145051d6

Pith citing papers

Observation 2eeb54c1-f742-4702-aacc-667cb7965203 · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:12.968388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:12.968388Z digest=sha256:e8fb2a62aca65d27a7cfddfb904b01e2698cb9f18fbdf31130ad94c6a7dd7fab