Pith. sign in

Paper Citation Record · LEDGER

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.07106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07106 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:12:22.931098Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:06.999018Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact14
  • verified fuzzy13
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f30be200-80c3-4379-a9f3-a4352cdede76 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.373546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:530fea8b106b73e661e99da98429e108f5975b2230bd1a3ae43b5f32e27c64b0

Observation 2a2659ef-4497-472f-947e-1bc1c23f2b93 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.383327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:937044bea8ce13d5749cd5a02da3fb09dcea0f9cb5f78c415acd0c91741a9520

Observation b718ad6d-f801-4db6-aea0-66706690e3f9 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.366423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:d539ef76e7d33a1eb38eda55781f90046689d100a6ee6cc2e5e2518c50049ae0

Observation 19d474a0-ae56-435c-92a8-b4c16eccf2b3 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:34:23.540456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:4431d6f3910c74193d935e8157b4c3f7808a30d41de5a017051138e4678a8f58

Observation 3adfb8e0-4d71-4e5f-b3a2-9fb6e0822618 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:17:58.977376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:0f8e20f977ed755aeab318763f9f9b426abc680e45b7e431646e3eaddf1b5864

Observation e5d4afc5-7a43-4bcf-9b08-e5f02efe6293 · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Llava-plus: Learning to use tools for creating multimodal agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.445108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:1f954a96c613d3efabc9e1fbb3e5dcddb3e6a57ccf6580b99197c683c86d716d

Observation 730654a4-c88c-4cf8-9781-d7fc512a27e9 · outbound

This paper cites Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.640305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:ddd083a4bf6458c7730be9c52ebc21c54d3ed170f4ed40ecaa3d581ffdc42906

Observation 772fc35f-fb3f-4c72-8320-ed114569a506 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Visual programming: Compositional visual reasoning without training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.452265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:5abe40a6d3fa140167ecba2a8317d54b3075a3b2b7439be7c2ce1a0cfe9597b5

Observation 049a28f5-82cd-4c0b-9b02-2f215d44fe27 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Vipergpt: Visual inference via python execution for reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.468839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:649421434ccb2c20c1623cfae9896296e6153d1c475dcaa04a98f1dda42f8349

Observation fd48fd69-35e0-4bb2-804c-a5b2753e6e72 · outbound

This paper cites Visual program distillation: Distilling tools and program- matic reasoning into vision-language models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Visual program distillation: Distilling tools and program- matic reasoning into vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.437297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:352c1dbceac9dd027a0113e63d22ffbf85f4d86c70b0511f05f5fc264c74ee30

Observation 29593fa8-184e-4cd6-a4a2-33cad8972be6 · outbound

This paper cites The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:55.088792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:94dd3fff3cfbe09b6eab249e433ca4d4525f1b47aeec07f9596079476dabd304

Observation 402a1fec-1c14-4069-919b-65d696137a6b · outbound

This paper cites Latent Visual Reasoning.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Latent Visual Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:d30c1b67ce25ec01e0bd5306c42d649cfb86e945d0100644925772a57b43bbb6

Observation 5da0f996-d261-4397-a636-c57c874ce105 · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Monet: Reasoning in latent visual space beyond images and language

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.550348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:af6e4cfef17aff368dc5c47622b7d39d5b84de1e6769ade10962d6e8de041726

Observation 1f98e667-4a78-42f7-afa9-c586a6b07fe4 · outbound

This paper cites Imagination Helps Visual Reasoning, But Not Yet in Latent Space.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Imagination Helps Visual Reasoning, But Not Yet in Latent Space

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:14.979939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:a89160711783278e1c9a2ca543a6dd4918d9b0e28cc4b9f1784a2d5b5b80df19

Observation 151ac1cc-7389-4b8c-9675-c8f91395a498 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:55efd0e6270a7ce5910a2d5a91f2c87f368188f7ecf75c8b30dc1cf746630382

Observation 3ea7456d-c685-4ce1-bd30-b3387d2be9ef · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.866178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:e343efa4ad7185d72c875a76e0115f5fbb46da53f1b3133ff7834b42fa6114b1

Observation 24f70aeb-2147-4bdd-98a5-b9771943db9d · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:29:05.896693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:926eb2215bcd51e4a031d03e068bd20ca88d3385acad5f163515cea90de53205

Observation e5910887-2f0f-4c79-8e06-dd73ae5c1bbc · outbound

This paper cites Codi: Com- pressing chain-of-thought into continuous space via self-distillation.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Codi: Com- pressing chain-of-thought into continuous space via self-distillation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.461386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:e97fa10ca6c1a2b2f11df0197de9184132326b63699f0bd716f1beac7f1b7c7c

Observation 4cbd606c-8f58-4820-8a63-acba136633d7 · outbound

This paper cites Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:59.524353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:60b1cd22dbd120d11be5094e03de472844caf946779dd246c460418689264dc8

Observation b3257744-8fbc-4791-9ae0-3a7d3ed086b0 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.653902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:bbfc99652eb8c3d849a81e98c19fd9091a5d8372c0dfe46faf304bb0440e9ff4

Observation ade04805-f270-47f2-9564-23bc3b7da21f · outbound

This paper cites Hudson and Christopher D.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Hudson and Christopher D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.388478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:6c51aa00a216bb6fcf561030f5c8390257464f0ceb5e138fe64e4af76e98bab6

Observation 17493145-aec6-408f-828c-c56646fb4f58 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.609673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:af7c5ca169fe0afb38fb5c30b584134e5c5950b65abfafb532e3ab4cbe0ffb1c

Observation 0ecea5ee-58eb-4429-967f-0d32834f2e75 · outbound

This paper cites Qwen2.5-vl technical report.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Qwen2.5-vl technical report

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.399164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:72a31dac77d9466be4beab0f5dd0d540e2cb17da77af79f5d302297055c6c956

Observation 4fd9394f-8fca-4044-86be-7736e1e946ca · outbound

This paper cites Qwen2.5-VL Technical Report.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Qwen2.5-VL Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.633185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:999e044672e564d3278ac4cf61a7f2c8d4657c5c28a180a3b7e26ff7c3c5f911

Observation a309c390-f6e4-4aca-8955-d4c92ae5057b · outbound

This paper cites Emogen: Emotional image content generation with text-to-image diffusion models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Emogen: Emotional image content generation with text-to-image diffusion models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:15:51.823830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:e1030b2942fac2be9d77fc203c24f464bea9fcb8139a507c347c804c91f4be83

Observation 53566c6c-b41a-4d84-b6ff-701b49d4a4b9 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.409051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:a75d11898e78fca148c6083111ea43da7d7c8111acea01874981bbd0f5d00408

Observation b2a722a3-2039-47e8-943d-ca67f1c46170 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.428672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:bbcb8a4417f684324d7d2b1fe18a37b0b1244aaf4168028d7e721261c509f26a

Observation ab48df36-fd47-46a8-af7d-1414478f447a · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Blink: Multimodal large language models can see but not perceive

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.420059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:b4962d0b3e3eaa80471704f0bd7df2e0c338f4d5487f0bddf9a01a69df5a27f5

Observation dde8b3f4-8951-41de-841d-81380f1b3855 · outbound

This paper cites Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:569073d0b2ef071aade64caaf3e1aa61384e5336a7b0e7112513e1ee9d899270

Pith citing papers

Observation 7a903fe6-ff2c-4282-b772-bc3acbc5e319 · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:06.999018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:06.999018Z digest=sha256:ebde82d6581d0dcd3256444f4c95947f8e08950deae5e17c288cb5ce0f9ca155