Pith. sign in

Paper Citation Record · LEDGER

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

As of 21 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2506.13888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13888 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:03.367306Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:33:36.634359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T21:41:16.983561Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71f73914-ffcc-4ad3-bd6f-45c80ff82150 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:53.663253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:53.663253Z digest=sha256:c5b94228368cf0862dfaef79e71563f57703253540ba8ea5cc9f63a8528ba4bc

Observation 5079a9e2-c7f8-4ffb-9c4a-243bfe007e95 · outbound

This paper cites OpenAI o1 System Card.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training OpenAI o1 System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:53.903503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:53.903503Z digest=sha256:83b1df63b1901f598ede92d2102329a679d2361c49d0a2135c4963288eb74f16

Observation cf97bcee-cb13-452b-a8e5-636f7213a605 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.034506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.034506Z digest=sha256:d8772b1568e320f72955b05265c3f35fa2835cebc734d572bc26f104007ea3da

Observation 81c9655d-a4f7-480e-9544-4e8fff1455be · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aligning large multimodal models with factually augmented rlhf

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:11.216046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:54.154868Z digest=sha256:b5d165d4a1273216d86c1064ea3040ae1cec1699ad86067898697fc926f8ee2c

Observation 5e70e07a-1c5a-468c-b233-6d55cccb9a60 · outbound

This paper cites Silkie: Preference distillation for large visual language models, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Silkie: Preference distillation for large visual language models, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.902201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:54.359009Z digest=sha256:3a3b20e0498340cc8718a977d4c58e927d678103d4c1ba8b3a93b22f64703821

Observation 92dc206b-2f27-4936-aa35-eeeed62849e3 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.489101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.489101Z digest=sha256:ba398b5ceb3cd90f5af82683c7c8162905d332019c71579e9e3f9cf654789c09

Observation 6e16397a-fb4e-43db-90b2-7bc3c5970f64 · outbound

This paper cites Rank analysis of incomplete block designs: I.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Rank analysis of incomplete block designs: I

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.584748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.584748Z digest=sha256:19e16504f22ef39097d1eed5aede2a331a4c25d4be42d297ddd727e09747bfae

Observation 19d226a1-faaa-4712-8c19-de7c679ac146 · outbound

This paper cites Strengthening multimodal large language model with bootstrapped preference optimization, 2024 a.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Strengthening multimodal large language model with bootstrapped preference optimization, 2024 a

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.724456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:54.714834Z digest=sha256:5814ac3024a3bbee539d1e9070650aab5e8bcc9fa206e20a9344b653373905bf

Observation 92bb1aed-1ce9-4191-ac09-1bfdbc1d0dee · outbound

This paper cites RLAIF-V : Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RLAIF-V : Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.836665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.836665Z digest=sha256:52c3edb60ae1381f788be1c368d1ce6b286dd0d2dc80205f8d45009274d2829b

Observation 8407a231-d41b-4fa5-9edd-13a16d08b29b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Direct preference optimization: Your language model is secretly a reward model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.525501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.007073Z digest=sha256:331fc3cc217a3a94fbd82b75e74267d69dd001b4cdff0a5d94895624ea6c083f

Observation 85f44ca7-2d2b-438f-8211-0fd8af44114f · outbound

This paper cites Self-rewarding language models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Self-rewarding language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.219552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.131233Z digest=sha256:4828649a83ed366b08951b362be3bab3eaef48e5a42175830a5e25cd886c1f53

Observation 897f665c-b0f7-4224-82fe-a734ff476882 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:09.840092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.313774Z digest=sha256:d8bdb3ab400dcb6a853b4bffedcf9be046be30a8cdc21caf8b33a0477d0f7dff

Observation 43432bc8-86d9-40e3-8f00-ea415d2bb00c · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:55.414768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:55.414768Z digest=sha256:ded3a4364fd5ca7425b5dca468bd363c28f38a6e28e902b793a3d57b0d55fd57

Observation aec22b80-3c18-4d55-b9d6-c941a1c09b7c · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Star: Bootstrapping reasoning with reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:09.535629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.542063Z digest=sha256:374b43da733c5505ac3adecc0879f440327245429f62244933b03debb9c63671

Observation ea707d35-2e93-4723-aeed-3f9c65a8d884 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:09.221417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.710576Z digest=sha256:eedd13c4ebe442b782ea243d0bee814d934f48a9a68741baf15d547022d8e431

Observation 9d3a3448-bd1f-4dd3-9b1f-e59ec1e1b857 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Rlhf workflow: From reward modeling to online rlhf

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:08.876205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.944239Z digest=sha256:ffc5504c0e9acbef8863d560b6ef080e47e3f06e5116e580f2e851b0a4380ade

Observation 335a85c3-ee6c-4059-8461-65bf81add525 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.124755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.124755Z digest=sha256:8c425d655b37c8c9009569843c78c5865bcdb3ee145bb9e3381d131a73e7a4d1

Observation 8ff8a13b-870f-4430-9788-4a5a244680df · outbound

This paper cites Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.251004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.251004Z digest=sha256:cb9d54273c28d252012bc3b6e5755d14b23a7e332c07afa0da57195f50740409

Observation 7d51a540-97a0-41d0-9202-02ef0d0b3ea4 · outbound

This paper cites AIDE: Agentically Improve Visual Language Model with Domain Experts.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training AIDE: Agentically Improve Visual Language Model with Domain Experts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:34:04.544547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:56.525727Z digest=sha256:0f9b0e7fe690df1a3f59c8974c6bd2a33904b96c3309f5ae7b42a1500d26a8ca

Observation 5f4fd0b5-7eee-4cc1-92ea-93643c245a82 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.665520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.665520Z digest=sha256:43d5a49ec0add633c681897a9af4cf3021f2aa7708ec78cb6b42d5773294c53b

Observation 4dc8e471-1f47-4447-9f08-d0e01cf6afef · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.770706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.770706Z digest=sha256:29e3314af871e0762945793e83ad97685e29afae4b5c64b7487108c983c4a592

Observation 17375a3c-39b5-4c5f-93b5-1113f5dbca4a · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.851632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.851632Z digest=sha256:c493e62efb9578d51d3ca41ef77a5cb11eb2faf0b9fabe3e6d05d71f9232f581

Observation 712e93bf-a4be-4b77-8f3b-32299be42706 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.965198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.965198Z digest=sha256:f21dc9b159fa4e6bb24d683df98f445267dc9672fb670db9b431568a925dcff2

Observation 97ee5a45-5a24-47aa-8ddd-27be7773c5e7 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.076467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.076467Z digest=sha256:1c27a395294ff9b202a1f302720549fa012e6be3eb6fad5bfed9140503f418ae

Observation 3af1a78a-4148-4762-8f6b-41e1dcaeab74 · outbound

This paper cites Chi, Quoc V.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Chi, Quoc V

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:08.485421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:57.145526Z digest=sha256:1a8dac0d3ca9bfa05253ea3f28a33fd570a8be1a821663a03347d30b40494e77

Observation 7300f65a-3223-421e-b689-8e8ecc2a5f18 · outbound

This paper cites Iterative reasoning preference optimization.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Iterative reasoning preference optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.337884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.337884Z digest=sha256:2a145afaf62eb156bc12b25f89c9a046839bb3de4146725bb87fb8c7fe47c5d3

Observation c92ff6c1-a64a-4505-b622-bf0e88232d6a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2023 a.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2023 a

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:08.231425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:57.515024Z digest=sha256:e5ddbc1f1ffb0c7fbb62b462e76ce8de46205690764fea799aae2e1a7fa5898a

Observation 2fda5150-1f12-4cfd-a2f9-b595027bccb9 · outbound

This paper cites Detectron2.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Detectron2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.605814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.605814Z digest=sha256:c3f08ac3c6468305a422a3fc6a9c6d13a11c10f71e9aa0369670a603a45b4c66

Observation 057fb7c3-9c81-49c3-aff1-b0695aace441 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.696410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.696410Z digest=sha256:2c727c5af9a0407ed544962980527c441b3e7de2b3f90a5b019a6ec68a70246f

Observation 9234ae8b-c117-484c-82ca-a7c707cd8634 · outbound

This paper cites Training language models to follow instructions with human feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Training language models to follow instructions with human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.824861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.824861Z digest=sha256:0584e0ece385775ee24f90fe60ce92d5ce6c7fd1ce56d0f3d6c6df402ee69ca1

Observation aeb1c785-afd8-4663-972e-468a808ab679 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.925619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.925619Z digest=sha256:6ab5d80c0478d6e6453fb4273c89d85dbfd5b899847e55e92189da04162800c9

Observation 16ab5fe5-b16c-4bf7-adaf-8a74df19abd2 · outbound

This paper cites Hello gpt-4o, 2024.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Hello gpt-4o, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.089473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.089473Z digest=sha256:9ab673355206f4291bcc8ec1f9c89e32017b70bcf86c51356b2d6d8c735fba07

Observation 3aea31ff-d467-4ea0-a473-3a36fc822df8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.204744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.204744Z digest=sha256:f7565a61211f1ac89c1f886cfecb9497d2848a4f484b115c301b3658a1e7f47b

Observation 4f756b18-5c6c-41e2-a79a-6dc72c1eeef7 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Gemini: A family of highly capable multimodal models, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.881904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:58.323913Z digest=sha256:f98b6f9c617a31300949e476b80a79996cb99e586defea724d76a9a03446f33c

Observation 1d6f752f-6ec1-4570-a9c9-00d766f05db2 · outbound

This paper cites Visual instruction tuning, 2023 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Visual instruction tuning, 2023 b

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.494189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.494189Z digest=sha256:83faa99b9de2d99ea511784efe2c6788c8d41a30a5dfacbf9090288a0d8151cc

Observation c78671b4-f52e-4dc0-941e-2810b163f8bf · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.598937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.598937Z digest=sha256:d499dcd1210fd707c6d8f57d94716c29568a189fbbd15b43c908a04d49dbe61c

Observation eab265e1-d15b-4880-8499-3ffa908e199d · outbound

This paper cites Llavar: Enhanced visual instruction tuning for text-rich image understanding, 2024 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Llavar: Enhanced visual instruction tuning for text-rich image understanding, 2024 b

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.546774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:58.678265Z digest=sha256:3aafd492e860aaed99ba7d1943bf48383e25778cc00b1ff77bbbfb30920f2671

Observation a0a66828-c668-484e-9de3-7be0787a80b6 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Sharegpt4v: Improving large multi-modal models with better captions, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.820250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.820250Z digest=sha256:1706494f3bf601081402a0b43fd1f60a7ed3f2fcaa8d1dcd9711ea1e7f780791

Observation 5f8b388e-d3b8-47ef-b393-3048f178eb6d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.895296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.895296Z digest=sha256:41ce6c68cafe6a3ebbe17dfcdb105a621ab650759fc702b3fbdbf0a2a3d569af

Observation 17ff4a32-4028-4ef3-be13-82cf82910833 · outbound

This paper cites Fast r-cnn.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Fast r-cnn

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:59.109451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:59.109451Z digest=sha256:00f489371e5ca78a7379b3fbbad9076bf85de1519b2b935ab38b1457729135b9

Observation 90a65f8b-979b-4f87-a34a-38dbe01c0037 · outbound

This paper cites G-detkd: towards general distillation framework for object detectors via contrastive and semantic-guided feature imitation.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training G-detkd: towards general distillation framework for object detectors via contrastive and semantic-guided feature imitation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.130882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.257399Z digest=sha256:a38fddb7c6388f15b72d50e5866046c9a3a3d44daa3023461f2e7554ae52a640

Observation 659a3a35-0423-48d6-bed4-51a35d4dcdb4 · outbound

This paper cites End-to-end object detection with transformers.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training End-to-end object detection with transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.837984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.466805Z digest=sha256:5a91e077d1e11362582f89f93e40acd2faec02844b3897d9dda8f3049b6545c8

Observation 06efcf38-7a97-4ae5-bb33-b0c38f04a6d6 · outbound

This paper cites Global-local path networks for monocular depth estimation with vertical cutdepth, 2022.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Global-local path networks for monocular depth estimation with vertical cutdepth, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.645786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.614342Z digest=sha256:a9c7f8d496e552b2bca8349cdd0300ed4c2a4d7aa84c6ec99a220ad0bb7c165a

Observation 32f16727-5973-4eb9-8a5b-ef19c1b28cf6 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data, 2024 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Depth anything: Unleashing the power of large-scale unlabeled data, 2024 b

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.408486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.733953Z digest=sha256:fcae0686401db0ea5656c19abd6f97cdfae6e1151fd8ebe908815c91d9b94ba9

Observation 703dd32b-d862-4d7e-af01-9adb8b00d82e · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:59.917204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:59.917204Z digest=sha256:8ccb16d792dbcedd3d5a3434a9c62256b74b32da546a3bd23d239e2ceb4082f0

Observation d1bed38f-3a46-46f3-bd26-4df047435b01 · outbound

This paper cites Detclipv3: Towards versatile generative open-vocabulary object detection, 2024.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Detclipv3: Towards versatile generative open-vocabulary object detection, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.072509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:34:00.069034Z digest=sha256:e62b1e5709fbd5d4b3cd2bcb4d4e795a105f28ac300f81eda8965afb29158804

Observation 9b653fa4-3b63-43c8-aa2b-e9166cfb4795 · outbound

This paper cites Let's verify step by step.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Let's verify step by step

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.209577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.209577Z digest=sha256:aca388d9c0c6e43fa24bcdad7395c77f99925f0c20505b635411bf7099189a70

Observation 1273b247-f85f-451a-b5ce-456753c43268 · outbound

This paper cites Entropy-regularized process reward model.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Entropy-regularized process reward model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.313903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.313903Z digest=sha256:4a95a0c743d6e8612a61b9f076b75c2a850e2d49e0d7423ab9d96187f997c8fb

Observation cfe08292-f8f5-4f7a-8ffb-56766c46e2ec · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.459775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.459775Z digest=sha256:19cfd2e8c6cd6b6e9247cc1fba84c18562cd8f317b0eff32d1789050f1c07238

Observation 175b940e-df21-424b-bd9e-dc08f0673f5b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.538427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.538427Z digest=sha256:c41829071f245fdcb20da043f0e7784f55e571077a9e34c95c3ba505ad119ba1

Observation e3a2a723-284e-4cfc-bbcb-7508e1d319c5 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.668832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.668832Z digest=sha256:2e5fb70650b9301a6714f0a10f9d2746b2b59fc59b32e917ca987d343018afc1

Observation 41150ab9-a1a5-455d-a39a-e2ed661c1a3d · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.772326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.772326Z digest=sha256:323fddf68c6fdc67dd6051824f2ee02f3bff008101600a56b2b7bb76561adeac

Observation 46d38e64-5e7f-4d8f-996f-544aef9cdf00 · outbound

This paper cites Reinforced self-training (rest) for language modeling.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Reinforced self-training (rest) for language modeling

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.689725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:34:00.898171Z digest=sha256:7a85244f5bea55f12e37413db006aab1e6923234da9ee9370eff48b4f0c9dd31

Observation 0f6a21d9-2b56-4249-b0d4-f50e735f049f · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models, 2024 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Self-play fine-tuning converts weak language models to strong language models, 2024 b

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.457636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:34:00.997086Z digest=sha256:0e1a52c07bd78e09640eba7cc766312e3859dfe79f7ed59045d19c1a986ce153

Observation 7b990001-6d13-4f8f-91b7-b7665f47f180 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Lora: Low-rank adaptation of large language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.254752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.254752Z digest=sha256:dd1fae61ffdbc4246e1a72e251ba5db47afdbed7ca8374041cf5d5735dc51496

Observation 4cf8f6cb-6683-4aa0-95ff-b22f40faa089 · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.366014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.366014Z digest=sha256:0f330746039c0527f601d2f1ad326575fbac385c4affc99279519a40e6203f84

Observation 9f24e489-e291-46e5-b7ad-b4a4f303a15c · outbound

This paper cites RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.131722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:34:01.535272Z digest=sha256:b0757b8265f5177b679463c7c62272911ff9ba92ac41bacb1a647f8940a525c7

Observation 68bb5824-4ba9-4039-b49d-1727dc6a549e · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.674745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.674745Z digest=sha256:910edc029dc6e08e622400dc8201d09b4fd6733ac4c0ccca7fc8999a90e26c98

Observation a725c543-89c3-4ae7-8a4a-e3e1d6d57199 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:04.935766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:34:01.784926Z digest=sha256:38d0aae99794f59baacce99bb1f164c6d4683f7cf3f9a439297eeb5e37e818d1

Observation 497b1613-5f52-448f-88a8-1645b6bba42e · outbound

This paper cites Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.910015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.910015Z digest=sha256:f54aca5bd2088c427f6852f864296736d1bc0a0d8ea324aed38bc83a6a401341

Observation 10b8cb74-dedb-4513-bbb4-17d9cd73fdbb · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.982398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.982398Z digest=sha256:9359697f94b58b8ef307602e03a8d1aa09c474c0da9eb12dcb22ffa85a9dca52

Observation 964dcb89-b750-40de-87cb-9f150da3de60 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.117511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.117511Z digest=sha256:4ebc9f9aa5a566e84119f5dce2555b5af1e23e369bf0baa53683b33048371b61

Observation baff011b-f8e4-49ec-ae43-3ddf46008b7a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.276771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.276771Z digest=sha256:cfb28720caae5ebffa9070da13d6a7bb4decd8eafb4cba6ca835057a2ab630b6

Observation 54846322-0400-41c2-bddf-a584ee6cdc9c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.392041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.392041Z digest=sha256:29079b93f08baf4d5223ce1f7d4a7bd54fe799d176e5864fcaf2d21f9c831dd1

Observation c4f294e2-d043-4c98-983d-e5fd7e3cea53 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.549287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.549287Z digest=sha256:361651ec431aaaff0b4b01f175da162eddde8e85d5b4fff4baf588cc15d0c048

Observation a18b81a1-3763-4337-ae71-4ba9b7ef5f7e · outbound

This paper cites The Llama 3 Herd of Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training The Llama 3 Herd of Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.661492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.661492Z digest=sha256:81574e93072965eaecc86b2a01873a97fa2eecbb5cacba99b8544cdd9cc99e78

Observation ced36032-cf5c-4953-9d69-a793bc53113a · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.816427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.816427Z digest=sha256:b3659b34957521acdc0510e42dddbfed1b595883442409812cc4f5890292ed17

Observation 210b0553-5bec-4d76-9353-5a32a79b2165 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.960724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.960724Z digest=sha256:912b2cee1b5fe2501548569da8811ab1c5ff625ed57180426a44b1ede33de159

Observation d45607f7-0eeb-4f4a-9271-0dce9d0007ae · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.076564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.076564Z digest=sha256:334945c276e0e620e020d67e726e80303f4c3aa77aded7a9844ec75a1560256e

Observation 456f258d-cdda-45e3-a431-c54fb8de5b8a · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.216795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.216795Z digest=sha256:48831accc89127270bcd58ac88491668fa919409e5fc6a8c6d2ab669603a6b21

Observation 5876a8e5-1de7-487e-a2b7-ebbc8f195b8b · outbound

This paper cites Qwen2.5 Technical Report.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Qwen2.5 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.367306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.367306Z digest=sha256:505861244427e83260ad604c3288591ef806660390c739bdc65fb5d267bcdd2e

Pith citing papers

Observation f598c56b-c547-4a9b-887c-e6ea53c849e7 · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:16.993491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:4cdaa7db6f3769a0db357cb58ae61ac633a0267aad749230fea7cf5bba7e50bf