Pith. sign in

Paper Citation Record · LEDGER

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01709 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87dc45e4-5c3a-4b44-b641-e9305154b74e · outbound

This paper cites MineDojo: Building open- ended embodied agents with internet-scale knowledge,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MineDojo: Building open- ended embodied agents with internet-scale knowledge,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.062789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.283714Z digest=sha256:1528ed4cf8f0ada435fe90c5738f3ee4543e78425f44b155ede4e0fedda5b4b3

Observation 0298819c-8cbf-4328-a7a1-ce84e788a605 · outbound

This paper cites Guiding long-horizon task and motion planning with vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Guiding long-horizon task and motion planning with vision language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.048813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.290888Z digest=sha256:a0f0e56c726a9f60171a5ffd969ffcb4f0db99de614b5abe3266d2a985f45e30

Observation 0e2cf3ad-b986-4786-b06d-431500771a15 · outbound

This paper cites RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.033839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.296266Z digest=sha256:7402d9a99cbaa68f186b86975ae32ff8763718b6c0615bc6d14db8086142dd2e

Observation e059d53f-7add-4df2-8310-8f46e39ad889 · outbound

This paper cites GPT-4o System Card.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o System Card

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.301651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.301651Z digest=sha256:9b2deb52781c24ca362fc9a36d0a334d30211a51f1a80f2d4a1481fcb7a0b67e

Observation 88cd2bc2-5f58-48c9-b623-af5589a23aea · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.306951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.306951Z digest=sha256:1846fa33ffbfb2a285caacc2ac79cd3acdd0b6887e0fbad25c257f6fbfc68961

Observation 2c2728c8-37d0-4036-a589-887823782f92 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.311855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.311855Z digest=sha256:8a2d17cbf06f4432f2685dc37de41cea5fe1c3e7913aef12dec322b4e357089d

Observation c7ba22f9-c844-40a4-a82c-974a8572ea26 · outbound

This paper cites SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.019844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.317415Z digest=sha256:165aae97d9ca278ab3552c55e524089c145d516a6f735e7a63274f7a8b99b9f0

Observation 7780fd84-6fd8-4081-a9f1-6dc0f4b9312d · outbound

This paper cites SpatialRGPT: Grounded spatial reasoning in vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded spatial reasoning in vision language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.981742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.321859Z digest=sha256:bdfa3c45468d7547d9241583f37b5958b53a6da5b2c901ee559665281814e0dd

Observation 0af0b425-a73e-4e18-bdec-9a14837c73c1 · outbound

This paper cites Depth Pro: Sharp monocular metric depth in less than a second,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Depth Pro: Sharp monocular metric depth in less than a second,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.966492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.326320Z digest=sha256:05d5dfdaf2f0c9efcec797e83cb2c4021d92c29a6caf752c2a601c79abea0834

Observation c095fd52-af47-432e-bba2-c8b4e30fe240 · outbound

This paper cites SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.951835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.330689Z digest=sha256:ebe8b8ac802e4c9a84e0b884acea7603df3488fdca2c98044792f46effa05f62

Observation d0a5471f-d938-4158-b714-03cfe772a3aa · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial reasoning with vision-language models in ego-centric multi-view scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.334788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.334788Z digest=sha256:87fbdcb767d20e9388545819b27e957d35c1f31c50bcccd165a0a04c3ca81622

Observation 61b5f3fd-bcbb-40fb-8227-3e6125283936 · outbound

This paper cites Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.935532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.339916Z digest=sha256:b869c513fdacd270b7d0b55dc2f46e63d078cc774e9c86404bf3b32c8addcaf7

Observation 379014f2-0657-4b13-bf6c-f11a96ad612e · outbound

This paper cites BLINK: Multimodal large language models can see but not perceive,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models BLINK: Multimodal large language models can see but not perceive,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.920062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.344222Z digest=sha256:a1e9810b5b09732ee6a7efa94af34c0a8d599359f20e32cef3d78f6558541103

Observation 92576b2f-e40d-428d-b5b6-4d2fece80fc6 · outbound

This paper cites Does spatial cognition emerge in frontier models?.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Does spatial cognition emerge in frontier models?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.904450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.348928Z digest=sha256:593da5df46562d3e0f9e80ba52fc5fc3413c03b4ee5031ecf5c2384a80aa41b5

Observation 10da9520-ed85-4023-899a-c9abdafac454 · outbound

This paper cites 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.886760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.353214Z digest=sha256:ff327b16a4d48c3dd4a48d6ebedb7d4a75636fbd346ae665c7ec6e65743fa4a9

Observation 52425481-1a17-43f5-a128-b77b882d059c · outbound

This paper cites Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.870793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.357336Z digest=sha256:d55a83b0fdf01a053a9680c26fb9ee510e8c6b846e7b23f661bb521ab08862c6

Observation a188f309-9bbc-4fe2-99cd-3c1766bfa726 · outbound

This paper cites Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.362399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.362399Z digest=sha256:4619cf14a646bdd6ae6b981c94aa71e8df4b4a6189617ac985beb08bdfcc9fe5

Observation f656a0d9-57bc-4b85-b574-913e22555982 · outbound

This paper cites Perspective- aware reasoning in vision-language models via mental imagery simula- tion,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Perspective- aware reasoning in vision-language models via mental imagery simula- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.854558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.367294Z digest=sha256:7aa68efb1a5bafb4f2c6ed45a38eca2350611ec03dfb787fc01a34f08a9bb090

Observation 68fe7054-dc93-4cc0-a620-c70ac1b07b1d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.838503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.371626Z digest=sha256:1eeeed9ae5d3779a60e1180be5b16cf5b7f64a2bcff0ef3184c96dc051e6d533

Observation 476a93d6-80f9-4514-bf07-ed793b2fd837 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.376141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.376141Z digest=sha256:83e170afedf279aa047a40b6a87db2b2eccf95520f959b14a0e6e16d0901ad86

Observation 7f30990f-4b6e-424d-8950-e472965b64d5 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.822322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.381375Z digest=sha256:2ab24941d85dcfe4a4c863e883e892607efeedcc0cbdbca6228fc466f4b51bcf

Observation 519062f1-9c41-454c-9c55-856aa0b2c012 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.385878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.385878Z digest=sha256:6ba4524e790cd72db63e238ce623659d7300f3a4c1765c9adcc615ed71b11ffe

Observation ca127aa8-a717-4cf9-9b82-6f0af81548ff · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.804991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.390647Z digest=sha256:115ee24b7e739f1ef033c4730b4e3e0bf7c5ecf994194c37e3c2ae60ac937236

Observation c9ef6415-a750-4044-bbf2-9d2b8d5436b6 · outbound

This paper cites SAM 3: Segment anything with concepts,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment anything with concepts,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.785581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.395556Z digest=sha256:70c1c2db86ed6a5e93a2782b8730afa06452f3b54ee107e85ffde52ab8e15244

Observation 5d7dcb07-d277-4052-902e-303b3de99b80 · outbound

This paper cites MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.768789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.404699Z digest=sha256:6fc81debd53a1cebffd0bf3a91ba761f3847ccf67067721da7bfc4381d8d0b8a

Observation 5ca7e6ad-3183-4a31-91e3-627813761065 · outbound

This paper cites Cubify anything: Scaling indoor 3D object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Cubify anything: Scaling indoor 3D object detection,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.752144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.409810Z digest=sha256:774297cf430d7fb482126b43149cffcdabc63b69c5b811ffae4112ae65c5e70d

Observation 50d942a0-c2a3-4ed8-81c3-544d2814d9e5 · outbound

This paper cites Visual spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual spatial reasoning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.733925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.414585Z digest=sha256:a0027836bfde311f07a3c1e04db1daddb1cbfae2a2d44079ef15cdfece3c6382

Observation 464a342f-43ad-40ed-87d0-2c1a875fdc33 · outbound

This paper cites SQA3D: Situated question answering in 3D scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SQA3D: Situated question answering in 3D scenes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.715872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.419014Z digest=sha256:3eb91d7b0985af9299b0948541c2393367fb18c23db0380e4cb1d4b93221653d

Observation 687b4756-2895-497f-a76b-87570acbe809 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.423033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.423033Z digest=sha256:06c2dead97c41da3a1a1c1024447bc5250b071da0b306a69fd17ae6d301c5a1e

Observation c9df14e3-ee49-4a38-a992-99480078df7e · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o mini: Advancing cost-efficient intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.699368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.428269Z digest=sha256:f4b0fee2ce13fa2b0305da96149f179f2f3a9169e04e5abaa7fe5fe0dadf69bd

Observation cb2cf448-7aa0-448b-b327-272763c575e6 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.433519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.433519Z digest=sha256:f80bcead00111c210a9e8c6ed4bf77b29e86fe23ef93af71c58a8d78a8d4da98

Observation 3a236af2-8546-427e-8fd4-cb0e601b11cf · outbound

This paper cites SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.683097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.438252Z digest=sha256:74b2c579797385a326d1eae36e08944bf25a09b0355b1d4b303c4d52e8c8222f

Observation 6ea41c6d-42eb-4ab2-9905-b6372b374b1b · outbound

This paper cites SpaceOm: Spatial reasoning with extended thinking traces,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceOm: Spatial reasoning with extended thinking traces,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.667452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.442696Z digest=sha256:50dd8a9324ef3778860094374d231528289b9ddbef530be959bae46a9708fb8e

Observation 4a941714-7708-4e66-a5ee-6ccbe85f1a95 · outbound

This paper cites Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.647404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T22:30:23.447065Z digest=sha256:da7c21132dbd1f7f36c32a55e461c922ba3cd6e3416b5dc0bc4ea7f852833489

Observation d991c080-7989-47ed-92cf-c07e2497874b · outbound

This paper cites SAM 3: Segment Anything with Concepts.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.400209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.400209Z digest=sha256:6365a19a5539dc71d74b5fc6276958a15288f54d18096838cc2bacf162d8c775

Pith citing papers

No inbound Pith citation observations are available.