Pith. sign in

Paper Citation Record · LEDGER

VLRM: Vision-Language Models act as Reward Models for Image Captioning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2404.01911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01911 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:15:04.420614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:12:34.682493Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 811f7398-0b2c-445b-b4d2-9ad1df5ff3ae · inbound

ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models cites this paper.

ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:15:04.420614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:15:04.420614Z digest=sha256:0f8794264e53b45e6f5c37ab9042a18895f3d9c8e9653f796e4e65861a2ef011

Observation 25197f98-ce85-42bd-b7fc-34da4153e12a · inbound

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models cites this paper.

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:18.007439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:05:09.254086Z digest=sha256:7df9f3d98e97a60ba93119fa2633831c3dc219f65be16595cde143db1b32d03b

Observation 51cd10fd-0a78-4d54-9f02-5d980671c7ef · inbound

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers cites this paper.

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.573580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T20:30:29.032353Z digest=sha256:f252655f20db87bf94e43a826b0606aaac3360884a8e1b7741ffc38f4cacfee8

Observation c02c1e64-9071-4be2-b368-8855798c18e6 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.569828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:8367b5abec4375ed6b0fd95ed10073246cb2cb02ad1e6f2454db2d5ad0c4501b

Observation 08a734bf-3648-48ef-8de3-8defe27e8707 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.684066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:9fcc62805501e1f5fb5ceec2de6a40f0953aef8f5654f6bdabaeaac2bf0483cc