Pith. sign in

Paper Citation Record · LEDGER

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2505.05464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05464 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:06.746182Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df9585c7-f145-4952-9ca8-d51d77c80e69 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.876272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:cd6e25cfa2c8335a4e6baa61a8f9db363a3f0a4b90f49a8e323008f8b711cd65

Observation 52166a51-173b-4065-baed-1b8849038595 · inbound

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models cites this paper.

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:50.323121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:50.323121Z digest=sha256:44d021bfd6da054510c39c59ed1736fae84fbb4ec218facbc8add39ae4bf9a0f

Observation c9568ce4-e1e1-418e-97cd-37470082b171 · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.819405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.819405Z digest=sha256:185a7450451f0c86443368583848dc7a5c02674356455638ebf9f0813bd3685a

Observation 18bcdaae-1d8b-42aa-9472-b489a4c8ac21 · inbound

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models cites this paper.

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:35.625880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:35.625880Z digest=sha256:ad94c01c433a7c62dc6fbbcc8d2e7bde04df29583365b91c9bcdda3a2dcdd931

Observation 3c15fbc6-41d7-49d8-9fe3-6b4c95398690 · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.779148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:f881bc77b62ce8c417ed2b61cae183161c0dd20f8fafa9ef44e40aae9986dd64

Observation 5e991e75-8394-4b31-8ba6-d69525255154 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:00.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:11a92c5e3413d9f10bc7ca0791a5cc7de16af17d2d9f501df66d7cc968c78355

Observation dd4b5312-f3f0-48b6-aeb6-f3645440beb6 · inbound

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails cites this paper.

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.487917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T23:03:01.151745Z digest=sha256:1bf107f521e0b7d96b7dcf7c96c84c0f91c899866c7718c1d2b48af069ba0222

Observation f9a56e1a-b4b3-4e7f-852a-9ff13eb396d9 · inbound

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning cites this paper.

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.623754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T15:45:26.891621Z digest=sha256:e5d1f3e64f061c72dc3a7f1e286fd52432e43264ba58570302dc375a24526a42

Observation 613f2f3a-19a1-4ec1-a37f-867336c1a40e · inbound

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models cites this paper.

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.348287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:10:12.974244Z digest=sha256:89140bd57d56fa7289f72357daabd35366bff55c0ca23cc71327c2b948b2ad4a

Observation e4417778-bbfb-4451-8ae9-3e4b983a85e3 · inbound

Recursive Vision Language Models for General Symbolic Reasoning cites this paper.

Recursive Vision Language Models for General Symbolic Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:11:24.638371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:11:24.638371Z digest=sha256:2d1e76370f9ea0a262e508343297cb1ed88cf8462043a637337a458adb5556a6

Observation 97c0e90a-481d-4511-b135-8327fc93146a · inbound

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging cites this paper.

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:06.746182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:15:06.746182Z digest=sha256:bcda6904613662887f562375cc5a5b87abb6ea6b2dc5f73b17260a6b04ae6c3b