Pith. sign in

Paper Citation Record · LEDGER

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01709 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87dc45e4-5c3a-4b44-b641-e9305154b74e · outbound

This paper cites MineDojo: Building open- ended embodied agents with internet-scale knowledge,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MineDojo: Building open- ended embodied agents with internet-scale knowledge,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.062789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.283714Z digest=sha256:99a1026efc4e1bab308f30b01c88a938f69bf5163a08cbfa2d955b42bcd735ef

Observation 0298819c-8cbf-4328-a7a1-ce84e788a605 · outbound

This paper cites Guiding long-horizon task and motion planning with vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Guiding long-horizon task and motion planning with vision language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.048813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.290888Z digest=sha256:3bf9c424df02ecb0f663daf0bf38c58546463da007ff77c6f5c452239e9c6d24

Observation 0e2cf3ad-b986-4786-b06d-431500771a15 · outbound

This paper cites RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.033839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.296266Z digest=sha256:6ac76280f2bac8e54e3307254a2200766e72234ad55330f50312c33cc467718c

Observation e059d53f-7add-4df2-8310-8f46e39ad889 · outbound

This paper cites GPT-4o System Card.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o System Card

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.301651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.301651Z digest=sha256:eeae2e55d785b6b91a820a3a6678116d7d4dc6d7c01a7b7b608a2d043324d0fc

Observation 88cd2bc2-5f58-48c9-b623-af5589a23aea · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.306951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.306951Z digest=sha256:9abf0260ec5cfe59b00ff4dda62c109a66d43d03ae4fe19c4a84006b7a29b9ea

Observation 2c2728c8-37d0-4036-a589-887823782f92 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.311855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.311855Z digest=sha256:13540b7a229aae2d8dbcba070c32d526934bd038b4e6964ec396df9805a89b31

Observation c7ba22f9-c844-40a4-a82c-974a8572ea26 · outbound

This paper cites SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.019844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.317415Z digest=sha256:839107464f81aa7a8af13ab7ab1599029849464e085174cb7580a0cb3fdfb6cb

Observation 7780fd84-6fd8-4081-a9f1-6dc0f4b9312d · outbound

This paper cites SpatialRGPT: Grounded spatial reasoning in vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded spatial reasoning in vision language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.981742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.321859Z digest=sha256:ca42afd71347ac6bbbc25767211433f899ed21dd302926747f0b826bfc4c6f64

Observation 0af0b425-a73e-4e18-bdec-9a14837c73c1 · outbound

This paper cites Depth Pro: Sharp monocular metric depth in less than a second,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Depth Pro: Sharp monocular metric depth in less than a second,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.966492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.326320Z digest=sha256:fc88f82d71462c167c70ad70f22ee3929be851dda0f0a56f0bf5a06ec70bf96f

Observation c095fd52-af47-432e-bba2-c8b4e30fe240 · outbound

This paper cites SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.951835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.330689Z digest=sha256:67db1262facc20b9f06cafe8c7001b46f29bb0f8e2bfadc423e86e772f0ee0bb

Observation d0a5471f-d938-4158-b714-03cfe772a3aa · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial reasoning with vision-language models in ego-centric multi-view scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.334788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.334788Z digest=sha256:3e057504531682011dbcafdb88b799eb3e793093481909061ae6f5f675788075

Observation 61b5f3fd-bcbb-40fb-8227-3e6125283936 · outbound

This paper cites Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.935532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.339916Z digest=sha256:de5d7cc7fae85faa096f84597a285ab8c7af3521b02550dabb1ff829e17ce888

Observation 379014f2-0657-4b13-bf6c-f11a96ad612e · outbound

This paper cites BLINK: Multimodal large language models can see but not perceive,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models BLINK: Multimodal large language models can see but not perceive,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.920062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.344222Z digest=sha256:13ef2680a11089f21ab067460ce2395ed0766dcd2d44445d741c811c5338be43

Observation 92576b2f-e40d-428d-b5b6-4d2fece80fc6 · outbound

This paper cites Does spatial cognition emerge in frontier models?.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Does spatial cognition emerge in frontier models?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.904450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.348928Z digest=sha256:eb453ca2889e01fe40d739d111df69c6e1cf98cc94a96fd0c0113e0eafe0330a

Observation 10da9520-ed85-4023-899a-c9abdafac454 · outbound

This paper cites 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.886760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.353214Z digest=sha256:c0725d9413647452d45d27f5fbc18ac4517ce6b1c96512c58e7a803cb6ea848c

Observation 52425481-1a17-43f5-a128-b77b882d059c · outbound

This paper cites Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.870793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.357336Z digest=sha256:39afbcc449431f439f1f093eb9730a4c00bbbc1aed80adbd0b1d28ebc901d409

Observation a188f309-9bbc-4fe2-99cd-3c1766bfa726 · outbound

This paper cites Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.362399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.362399Z digest=sha256:a5bdfc255e4e54836d5b8cadd340232d5a5eb502d41c9378f3357ad7c00cf38f

Observation f656a0d9-57bc-4b85-b574-913e22555982 · outbound

This paper cites Perspective- aware reasoning in vision-language models via mental imagery simula- tion,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Perspective- aware reasoning in vision-language models via mental imagery simula- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.854558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.367294Z digest=sha256:eb3bff554d7ae4c3548e07c586001163edb3469cd57db973d4fba82d7c1d69fa

Observation 68fe7054-dc93-4cc0-a620-c70ac1b07b1d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.838503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.371626Z digest=sha256:fd67a9e5f9d0b04767b930b8fe74fb2a2ecbc0775a897d7299cbba131f39d220

Observation 476a93d6-80f9-4514-bf07-ed793b2fd837 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.376141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.376141Z digest=sha256:f7e079ebb5cce2d337e005e353dc754f310230239e5a3c681ad143a47c656ff1

Observation 7f30990f-4b6e-424d-8950-e472965b64d5 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.822322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.381375Z digest=sha256:177c859140a46bddefd16afa353d9df4fcda81318c73a456362e9e5ea8430cb0

Observation 519062f1-9c41-454c-9c55-856aa0b2c012 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.385878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.385878Z digest=sha256:4bf979984930c23941295f1b5a813e1f1478ef6f7a5e734d6d95337624bd8192

Observation ca127aa8-a717-4cf9-9b82-6f0af81548ff · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.804991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.390647Z digest=sha256:8890e0a8a90d07174cff7b33d5f283c05d6810824366b97cce61138679415be7

Observation c9ef6415-a750-4044-bbf2-9d2b8d5436b6 · outbound

This paper cites SAM 3: Segment anything with concepts,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment anything with concepts,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.785581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.395556Z digest=sha256:5a48867aa50c1466538beca8acf38c51b5c6b6a5b88bdcfbe86a73acc1f0137a

Observation 5d7dcb07-d277-4052-902e-303b3de99b80 · outbound

This paper cites MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.768789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.404699Z digest=sha256:5b6b4dfb1359f23445c887c6e91df59df95e18f9ee7423799c00c858b923a152

Observation 5ca7e6ad-3183-4a31-91e3-627813761065 · outbound

This paper cites Cubify anything: Scaling indoor 3D object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Cubify anything: Scaling indoor 3D object detection,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.752144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.409810Z digest=sha256:11c42b2217b0c89be7f612d844721dc82d8d9f6aa32e7155bd482fc645908345

Observation 50d942a0-c2a3-4ed8-81c3-544d2814d9e5 · outbound

This paper cites Visual spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual spatial reasoning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.733925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.414585Z digest=sha256:391d679f7ee11f4954b457029388662b79faa97bf016f39137322675ef25fdd1

Observation 464a342f-43ad-40ed-87d0-2c1a875fdc33 · outbound

This paper cites SQA3D: Situated question answering in 3D scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SQA3D: Situated question answering in 3D scenes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.715872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.419014Z digest=sha256:9785dd249fbe372095adc75f29a47e3c93d9e7a7366549808637a1fe36d5dd4b

Observation 687b4756-2895-497f-a76b-87570acbe809 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.423033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.423033Z digest=sha256:b1de41ae914e94779fe850be67623a5b7c9e4531ffefde9bb66e0e0d1ba71e1f

Observation c9df14e3-ee49-4a38-a992-99480078df7e · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o mini: Advancing cost-efficient intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.699368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.428269Z digest=sha256:2213a5a2349188e92afe221732d2be7c01f4981c8b058ff149c4028f10e173f8

Observation cb2cf448-7aa0-448b-b327-272763c575e6 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.433519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.433519Z digest=sha256:e9a933730be1580ae3f224ceab5de4220dceccda045f9bea8f0d4e60ed853c92

Observation 3a236af2-8546-427e-8fd4-cb0e601b11cf · outbound

This paper cites SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.683097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.438252Z digest=sha256:5affca88d805a2fffc66e683764ceb62495157cd704693d0be8cbe298dae5a9e

Observation 6ea41c6d-42eb-4ab2-9905-b6372b374b1b · outbound

This paper cites SpaceOm: Spatial reasoning with extended thinking traces,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceOm: Spatial reasoning with extended thinking traces,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.667452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.442696Z digest=sha256:144de28dfbd4b00df6124f34c12919250917df9e37d448d1cbac978b8f34cfef

Observation 4a941714-7708-4e66-a5ee-6ccbe85f1a95 · outbound

This paper cites Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.647404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:30:23.447065Z digest=sha256:e5498089329f47282def57f135a489c385cf4ddfa3bb9549a1b7af801c54e8e7

Observation d991c080-7989-47ed-92cf-c07e2497874b · outbound

This paper cites SAM 3: Segment Anything with Concepts.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.400209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.400209Z digest=sha256:b6b4c2b62caabefbd3810b67f0bfab3ab678a13be487a769a81710e1a7efcfb0

Pith citing papers

No inbound Pith citation observations are available.