Pith. sign in

Paper Citation Record · LEDGER

Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2408.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.00754 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:52.468598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:27:43.983079Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9f2284d1-ed7c-4516-aa29-aaf7658178a5 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:43.990119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:b63ae51f5ad8401ad321536d0d9a8a8c60166fc64e87836e84528b5497d995da

Observation b48a3527-234a-4dac-87a2-444f5473da37 · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.468598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.468598Z digest=sha256:df7ae273a8665a98dec2e341bed0324ee726f39aafe5b0e2ee07963343f54114

Observation 68399f19-443a-4216-9391-869add710de0 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:18.039306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:18.039306Z digest=sha256:c8f64dd5ca77d78abe0073791582014f4b218c3b0996db1e1667aca9f8c44f55

Observation 06ca9977-97ca-4c6e-a229-b1cf9a2edc0a · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.740878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.740878Z digest=sha256:4520fecd0d4d11be9d6a208966e627a70630b66f0db4d093768f6e5fa3015934

Observation 1adc13ec-88b0-4da0-9873-fc780c13da75 · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:28.115899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:28.115899Z digest=sha256:4b4168f5973670556c648b77421938404991f521b5d2ba1a01c9b42b89dafa7b

Observation ed1433d7-667b-469a-9390-b35f87b0fbe0 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.084749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.084749Z digest=sha256:c62ef3db38a6a0b4ed8f98effa59cea7ab023d6d8cd11ed91fc530c4f1b08e8f

Observation 31c2b4e6-2685-4110-bcb6-df4e1146e7c6 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.045708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:9cb841286ba2a44de233c8aa8c79f56f14e0e0f31bc360986d15eca2e7087fd1

Observation 1c0cf0b0-9d36-4abf-b355-42620e1141c5 · inbound

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models cites this paper.

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:50.003291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:50.003291Z digest=sha256:3b36d2c7f361dbac17a3531d22b3e9470dac80c58807fb74adcb99d55d6bf488