Pith. sign in

Paper Citation Record · LEDGER

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2412.03704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03704 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.089127Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T13:24:40.386203Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d5603a6b-9393-46dd-8c6e-5c5d0a5c7a09 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.089127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.089127Z digest=sha256:06996066220d66689aee4444de538e774e8ef6b62cfdda68481acd619fa0b253

Observation b69cd9cc-54b3-45bf-9283-819838cacd21 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:06.014106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:06.014106Z digest=sha256:f58b4d883ea41f01255e128080226a9123bbb9f5c78d3fc08f9c588985b54457

Observation 246c9b34-c7f0-4ca4-bee2-4043cdf08d90 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:03.181743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:03.181743Z digest=sha256:173fbe734aff654aede2ea133fa2f1cdf4fca24703bdc26af9648251a29edaca

Observation b54b5691-eafa-49dc-87b5-550f54f6c437 · inbound

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding cites this paper.

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:18.661384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:18.661384Z digest=sha256:aec00ce50d2084ee6335a3e9183fb9cbc2b544a54d9090f19071cf157cffc24a

Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.741673Z digest=sha256:9dd5af1e5e604f308ac4ff6c6d1f72c492cac0200d2e7b7e0c2b10cabc4c5c7f

Observation 6add957d-08f8-48a0-9bdf-d3f310ccc30a · inbound

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning cites this paper.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.096140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.096140Z digest=sha256:8ae3d643326da1b1f1ddf851395e76f0c5e2914dbb766bb2a03c6ca840d22be2

Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:24:40.393917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.858355Z digest=sha256:a389b9c2c8ca5c2fa9ee413e2b5dc41572c22c769c6c0379cd58db55e8765da9