Pith. sign in

Paper Citation Record · LEDGER

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2505.05464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05464 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:01:02.057343Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00e8cf36-898b-4728-a796-78f4cd00c571 · inbound

Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic cites this paper.

Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T17:01:02.057343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:01:02.057343Z digest=sha256:a6cba49fc3f5ec6ed3bb2e147946ee82cf8df95bdc1cb23d2d6fe95869bc9568

Observation df9585c7-f145-4952-9ca8-d51d77c80e69 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.876272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:4f7fe4fe2630b3537c34f5b41b0c66c606799cf3b20d62ef3a5bdd118e77dcc7

Observation 52166a51-173b-4065-baed-1b8849038595 · inbound

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models cites this paper.

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:50.323121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:50.323121Z digest=sha256:9555c8cc54f4b24c15cd32cc2d83021255eb57c2758d743da65665f99aa418df

Observation c9568ce4-e1e1-418e-97cd-37470082b171 · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.819405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.819405Z digest=sha256:5cc7d5e22e458580bae8bb9db389af73c459524888f61a3ebeec1e7c6bccbd1c

Observation 18bcdaae-1d8b-42aa-9472-b489a4c8ac21 · inbound

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models cites this paper.

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:35.625880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:35.625880Z digest=sha256:c0f62d059648715dbdcab90380535db3ea5d831865ed18258af84912e73f799a

Observation 3c15fbc6-41d7-49d8-9fe3-6b4c95398690 · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.779148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:5719b4aaf00526b0af580f9443952f27b1a882f4b3685e43621a89039069599e

Observation 5e991e75-8394-4b31-8ba6-d69525255154 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:00.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:c56b491eae9ef2844ad82f6a823b0a2e69a9e02611c98e41dcf4541d249cea74

Observation dd4b5312-f3f0-48b6-aeb6-f3645440beb6 · inbound

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails cites this paper.

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.487917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T23:03:01.151745Z digest=sha256:152dc420d6b29f4452dd73b030ea69fed44d08ac55f3548d12d8df2a5d982cbc

Observation f9a56e1a-b4b3-4e7f-852a-9ff13eb396d9 · inbound

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning cites this paper.

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.623754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T15:45:26.891621Z digest=sha256:8c6d7ced2cf39c88d55f88497b2dc72cc0fed048632af03d0f258f9e684407b2

Observation 613f2f3a-19a1-4ec1-a37f-867336c1a40e · inbound

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models cites this paper.

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.348287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:10:12.974244Z digest=sha256:0e2c9a5fe15af65de06a2a85c5ad594eb93b5a9a6f32072cb7c102cf447574ba

Observation e4417778-bbfb-4451-8ae9-3e4b983a85e3 · inbound

Recursive Vision Language Models for General Symbolic Reasoning cites this paper.

Recursive Vision Language Models for General Symbolic Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:11:24.638371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:11:24.638371Z digest=sha256:a0f86a9cf80848bb0ed05468a54dfa014cb61c1e9940824558d3244fa371f67c

Observation 97c0e90a-481d-4511-b135-8327fc93146a · inbound

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging cites this paper.

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:06.746182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:15:06.746182Z digest=sha256:31edaab8e4d9510eae21bc2c9a3aeea395c0d57b3c43893dc6ec804301c45400