Pith. sign in

Paper Citation Record · LEDGER

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

As of 10 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2606.26535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.26535 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T05:31:17.461916Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5d0d744-97d9-4ced-831d-bf813336622c · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.241764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:256e3380b46cc4431ad61cbe1035dd32d9fe4248fc93857c43caf067b70a48e0

Observation 00c60199-c9eb-4423-808b-4bc8cff4f204 · outbound

This paper cites Qwen3-VL Technical Report.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:09:50.247909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:03e37c0b5ab11d80b729df5321264b6d9b5aaef5bd72efed1f2f8c2429ab8a7e

Observation 1a1779e2-c1e0-4fd4-ae3e-09ef5234f6ee · outbound

This paper cites Qwen2.5-VL Technical Report.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:09:50.229185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:7e1ad28f3ee6303161e94aca551cbce3587da39f7cc329d4a89d741b2ca9218f

Observation 83c946b8-c033-48d8-acb3-1d4beb3f441c · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T05:31:17.461916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:fb9363489c96e10083815fca74a6203b6932e9b2662acb891073199d3288f4c4

Observation 833a873c-693f-46bd-8890-ccd010bb8025 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.250608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:6eb9d09f550f19c01e51ffeb057430be4489e3fe5b5a232f2856bd426b0494ec

Observation 81e5ab5f-8e76-4c62-a060-bc47dc6e5a5a · outbound

This paper cites an unresolved cited work.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T05:31:17.461916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:df6b035ea2076d3981bbf246167c2111f497b4552ba859fbcd77aee359a5116f

Observation 01469653-fd64-47f6-bc32-e6f8c0ffde20 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:50.235779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:6a84434446150536c15d8d06f0b10ed33b6cb11b876866e8e41d68f0f263f2b6

Observation ccafd9dc-74d1-4d7e-a53f-23e3a632f9a9 · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:50.244898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:487bf76667e18cc5f48859b242225b0237a18d91b69c96fbd7e0f4680a215c1d

Observation 22096ffd-1387-47fc-a43e-05e21824912b · outbound

This paper cites OpenAI GPT-5 System Card.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP OpenAI GPT-5 System Card

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.259228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:24b28c53cf70c932d562bc1fcae6d713630bc3fcea8f92fffe458a13cd1c5992

Observation f5aa600a-80eb-48c8-ab3b-23a25521e6bd · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:09:50.232626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:2a31d7e597d03f00c571aa81f9f855c14026bf5975f8ebeaa15a62e5d466f108

Observation 0a00fa9a-1f7f-44d1-acaa-d6f3ced63d87 · outbound

This paper cites Advances in neural information processing systems35, 24824–24837 (2022) 30 Z.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Advances in neural information processing systems35, 24824–24837 (2022) 30 Z

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T05:31:17.461916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:8b4eda7aa7c63fa1bcf6f5a1664589038fdbb94fae5282c0e6b73ccf435f8c1a

Observation 7f09df85-b641-47e3-928b-7ce30b575576 · outbound

This paper cites In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T05:31:17.461916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:068921b87e702842c7ce9f5567693832ab069d1d25c9474de22acf3459742436

Observation 4928d0d7-7c39-44f7-86ea-0843df5c6cd8 · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Cambrian-S: Towards Spatial Supersensing in Video

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.253799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:6435dcd8c47064ecf62fa66198d8a85681181598b29f41d01cfc3ba2f99c3fec

Observation 7e10ce79-1938-4d77-a9e0-0639533f8f69 · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.238739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:5350adb180cb5982652eb23e9febd5e8f3f13ce1ebea9028126743c132003392

Observation ac3f3fc5-c4c5-4ca9-8737-f4f99e85825c · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T05:31:17.461916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:38e57d1e42ba826da9d7b198a8f8479fbbd3317fcce9efb9068921d4968132d2

Observation f59c9272-3614-4f48-a34c-532ae77ff0de · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.256764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:15cd6f468a0314fdb762ba7bdb9ce549f8457b2ea1e16af78d045a5092e3b20e

Pith citing papers

No inbound Pith citation observations are available.