Pith. sign in

Paper Citation Record · LEDGER

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 16 inbound Pith citation observations for arXiv:2505.22019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22019 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:24:54.668238Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:23:34.984840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.135392Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69531200-78c5-409b-9075-705dee559482 · outbound

This paper cites write newline.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:48.648661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:48.648661Z digest=sha256:87b1d5c934ace72c59c41045e1c986b87d6580c83340f14db417aa3c37475ab3

Observation e8ccdf16-eff3-4b06-8bc0-390b31df102e · outbound

This paper cites Qwen2.5-VL Technical Report.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:48.786792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:48.786792Z digest=sha256:0173b44d5c3504460e7316c4d03b575732a92527a4360f03532510635664befb

Observation 268315af-74ce-4a4d-8877-217f37f1bfc5 · outbound

This paper cites Benchmarking large language models in retrieval-augmented generation.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Benchmarking large language models in retrieval-augmented generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.946878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:48.906125Z digest=sha256:3decbcf65024a9a971b334de8d4bad8efaad54178017e28ef18a3c5c987970d7

Observation 5baac46d-a7d7-4111-b649-57ccaf6cddc8 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.741210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:48.998331Z digest=sha256:cedfb38b9d97aeadae5c434a0dce51cf47bc828a2b9df7c4943e4c63b8a649e7

Observation cb69defb-b57e-41c1-bb83-ccf06e98986f · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.147441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.147441Z digest=sha256:cb74d578a6061f8027f42755ec05d409adcfc74367daac3c926e17814f8d6cc5

Observation 0af78db5-411e-4aa0-afe7-7ceefecb0d9d · outbound

This paper cites Mindsearch: Mimicking human minds elicits deep ai searcher.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Mindsearch: Mimicking human minds elicits deep ai searcher

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.352425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.352425Z digest=sha256:c124e7f8cf522b29d05dc383357b2fa6fe8ab6f7a58d6d429c2f6d29bd3b6ac8

Observation 762efeb2-6d0c-4bec-97a0-d036d833bbeb · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.490774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.490774Z digest=sha256:a89c4029973e3ec4f1542fb054f9b20feb3552ea1edaa8d630314f0fa42a14a1

Observation e9790918-9649-4cc5-ab71-331c48752194 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.709222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.709222Z digest=sha256:a0adce207cd59e842e84ab6398280c7ee24b39e4d410a245ad688bdc6b9e3027

Observation d32f2860-0616-4982-b3a8-d69e8cd7924b · outbound

This paper cites M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.927205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.927205Z digest=sha256:e99c7a9816c04480eeb30895d3f5713dce3d3fcad9d46f90876bec96e0ec35f5

Observation b8925cca-2976-4e05-b1cd-19f0c477861c · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.065723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.065723Z digest=sha256:8e44c55a20819a6ee1c23db8d11490321faed4ffedf8ca224d1f7ffd17b9b262

Observation acb7fd87-8eb6-47a1-bd8b-4ed82d8b5d9e · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Colpali: Efficient document retrieval with vision language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.538027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:50.220936Z digest=sha256:213a54d6430fb626df6fbd777124d5b4a8d2b82737b9632320c981880197a915

Observation b48724e1-3d34-4973-8f85-220a4d7e13b9 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.367366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.367366Z digest=sha256:729d88eedc257b5d92173c3de3ca41c2becb2495a3ac8c6d7929f2c275006ff2

Observation 07fc34a4-9d1a-4da1-a76d-6a940277e99b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.460221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.460221Z digest=sha256:96ad96cd13331f4af3e124f1de7a4cd9ca39f25e02f50538835cd148706e2a90

Observation 7ff22933-ee18-42ad-9a6e-3307aa72aa7c · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.559672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.559672Z digest=sha256:fdefa661dde70c46572bfd4f8a9ab69aaf914b5f1ddf18fa0c38e740dec2be96

Observation 00eeeb99-9d5a-4aeb-9c23-ed590f381a76 · outbound

This paper cites OpenAI o1 System Card.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.683709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.683709Z digest=sha256:10091a87becaab667217eb9fa7d6a6734c0c8048361608869718e2c6518cdf5a

Observation 00e16907-51c7-4867-a3b2-3f372a8b503a · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.840704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.840704Z digest=sha256:aaa862ee704e8720ed55e6819c52c1d035ceab577947c77fb5d42ea38ff007da

Observation f7548be0-6977-4d9b-8ed3-e540525ea30f · outbound

This paper cites DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.943846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.943846Z digest=sha256:e150742a6541ec945b9bb97a2e95c1d04af3c827aec4b5389b03c52d532845d8

Observation 0914aadd-44d0-486a-a3d4-d761414d506c · outbound

This paper cites Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.082163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.082163Z digest=sha256:6f201c323e64ec6b1fa3014b1b2988426e554804814ee530ea409bf2bb68b84f

Observation 9ead5704-8345-4c87-852f-5f9e97d29a78 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.217474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.217474Z digest=sha256:bc48fdf79ba31e8aef30c4fcf55ec8e5995c12b23879674509bbb2e426d0dbfa

Observation 8e205582-82cd-4269-8724-f25b68dd7ad2 · outbound

This paper cites Reinforcement learning: A survey.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Reinforcement learning: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.351090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:51.389300Z digest=sha256:b916c038b14229ae84a7f299a41607705c039cac6bf2ef8ea7b6e50aa0bc804a

Observation 6d2d3ac9-3765-48a5-8255-04acfd34cbbf · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.601363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.601363Z digest=sha256:24f3e5d1064f48c7c97a11c9dace7d955a4dcd086d49eaffadd8977064b52ab1

Observation c2d12a3a-869d-4955-84ef-bb3ddd92a9cb · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.700223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.700223Z digest=sha256:7968c19af799c708d7a1b52e1c12c30129414710551ef559ee239a82854cbb57

Observation b0219256-bd67-4df6-9a69-389dde6c5c69 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.824122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.824122Z digest=sha256:e6ab234a95d0c24b30fc55a3243f0acd61a23ba0cc80f23ec77f3293808ac8b8

Observation 11459a76-2481-48f2-b128-0ef82e496bae · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.884343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.884343Z digest=sha256:858c98bd66d51b26ab13c091d1aadfb79d7700539f16c66c183a4af55298c061

Observation 8b9ceca8-29b6-4506-b59b-b18646e7e1e1 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.963166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.963166Z digest=sha256:29601d751f9ef200777d0085a6292e591c4dc0e5da0be2f16d9332ab5ad21d5a

Observation e8acbfd5-bd5a-4648-b7b9-5e14039ded2f · outbound

This paper cites Improved baselines with visual instruction tuning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Improved baselines with visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.036227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.036227Z digest=sha256:389c70323163c3d23c14b0011602bd1fed3d7be59d227c1ab0fa4336dbc7d1c6

Observation e0f17fc1-8a65-485b-ae63-665bbe9a6844 · outbound

This paper cites LlamaIndex , 11 2022.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning LlamaIndex , 11 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.146606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.146606Z digest=sha256:d77802c92bd18c775a1bf9e40f8e2653531ec38082526cf4a35d96fffad6bbbe

Observation 02dc48b8-0155-4e5a-a008-61cb42d6d12e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.248859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.248859Z digest=sha256:3a83d89492d91767bf477d771782e05617f0d9d118fc8a1180ac8ea0d8bb270e

Observation c0ed3f2c-aa1f-435a-90bb-bee0c89944d7 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.343132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.343132Z digest=sha256:f9231d124cb5d2763ec709b865da54d67e0a5b55e1a42a7d8483515b36f6ca95

Observation d634e7bc-32f8-43f8-a76e-bec59d2d6849 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.421805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.421805Z digest=sha256:8463991565a8daced9cad6f660de76d498ce9ef9dc1356b0a50bf0de7155bff8

Observation 2c2ea2cf-8642-469e-924b-799c5c19984d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.518896Z digest=sha256:ca64d5e315a082da3f37844f12a64dc67f7f797c0015d7d40b2159e0c0457b96

Observation 4c8c0d4e-b367-4663-93eb-1233a01ea3dd · outbound

This paper cites NV-Retriever: Improving text embedding models with effective hard-negative mining.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning NV-Retriever: Improving text embedding models with effective hard-negative mining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.610720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.610720Z digest=sha256:8e0bc4d068498174888d6999be17aed496a4a2ff9d281557d2c0fb1e32355430

Observation 4dd90c39-b902-4ffc-b08b-f7206b3b0c2e · outbound

This paper cites Hello gpt-4o.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Hello gpt-4o

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.679894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.679894Z digest=sha256:fa741c32bf368adb882845e7cc8cc19dfada95300e92da34d76a9ae836d05d5f

Observation fcafe132-9e4f-4759-bc04-037d07ee81cb · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.763882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.763882Z digest=sha256:4e3a6294f04af3d66d4c320c2771d38301ee199a61c368023bc2c3f1f7041f90

Observation d251e266-f212-41fb-bcc0-cba2aec026cf · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.861519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.861519Z digest=sha256:dcca20eba47c92dd80c8afa2661e7cb5b48dc056dad2a377e01ecde9dd17a68b

Observation f33cf807-ebaa-485e-92ef-8f9286a003c8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.950386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.950386Z digest=sha256:9abcb185e224a5ae00dc77b4baeaa0c5a79d825a74294c1ecb51cf96d6a2a2f5

Observation 65ebfb42-2e5d-49bc-9dfe-29506b85d40c · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.100544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.041498Z digest=sha256:2021b7d25bddaa1fc2269f6d7e554b911d40cfe04e5f4dbe73a020133686abd4

Observation d74a9cfd-9ee9-48f7-b84d-946c85862326 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.171877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.171877Z digest=sha256:d499aa0548b11596a341fc40a1c1c03c63b8051f1fc39da42b1e8856f9a8becf

Observation b86812e1-643b-4b23-82d6-b42c127aa112 · outbound

This paper cites Reinforcement learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.809640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.264735Z digest=sha256:86fea2a3ec31d7eb8b32d866dcd0b40f53625ebeee59f8d2b1abfb3d562f6d11

Observation ebfbe6e8-0a78-475e-9803-6fd86710633c · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Slidevqa: A dataset for document visual question answering on multiple images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.604920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.364738Z digest=sha256:2c0cc69156a8425924dce420214755d9d5b4c7a8b9dabf410cb825665875b3a0

Observation 056425ae-5ace-4508-a89e-5f31bb5c7e35 · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.442043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.442043Z digest=sha256:ab8b942ec7203a0b119abc2937da6286594154b03674035adf13bed217f50c8c

Observation 6ebaea24-4713-4d9a-a30b-d63e2cd79fc3 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.487266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.487266Z digest=sha256:3818a3151af8d82721dd239a540a7bea7c3769a33898b89a32938f7053c0b276

Observation 5bd4bdc1-02f8-4d2f-98c2-88b1d53c7328 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.586633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.586633Z digest=sha256:2f2f433d4581c2a44ee980769ca40fdfb368eae9e802d697f042b26bff6f4278

Observation f15666ee-bdbc-4691-9a14-5246d2f099a3 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.713826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.713826Z digest=sha256:2d5fd10de0631c17c0de9f51683a35cdcf4b0deb280e7d7b742d0cccdd66e8b7

Observation e328ef64-c72b-47c8-8eda-6451b5932802 · outbound

This paper cites Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.810076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.810076Z digest=sha256:564a58e2839590a7fb6cc8ea1591cd051b42d53c4f8f3ef6d42e4b8e71002dfe

Observation 04b2de5d-2172-4ee4-baca-17c71b3a49fa · outbound

This paper cites Rule: Reliable multimodal rag for factuality in medical vision language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Rule: Reliable multimodal rag for factuality in medical vision language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.380239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.914447Z digest=sha256:aeef3a6de71443ee4d0f5fef37a51e853d4c081de4cc3d2f3f49692d815636d6

Observation 3ec533b4-47bb-425c-bc96-4f0808b4965f · outbound

This paper cites Qwen2.5 Technical Report.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.021458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.021458Z digest=sha256:b99fb80fb75d9ba47573e575347b310eb4b4852983483afcf4afa435a618d18d

Observation 161fa363-a68a-42c3-8582-5de71460e329 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning React: Synergizing reasoning and acting in language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.132125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.132125Z digest=sha256:9585936a78e0172bbd0ffe464d36fc643223655d75009b4730c67bb7e79ed833

Observation 512576cf-8bdd-4828-88db-3d113f7ce37f · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.235680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.235680Z digest=sha256:19de8527a5e96d79788195ef17a8e076de45ce62ab828f5c780ee06bb5437392

Observation 59e799c9-745e-4f80-8775-79913f385c70 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.345474Z digest=sha256:a374b3a8ce0b57479bf73129bad5141164a8a6e9f55b6f4968410ab141104060

Observation bdcf5e24-caf3-41f1-8888-177af150b37d · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.427481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.427481Z digest=sha256:78630f8b1448fe3179138755d7ebebe75d4b994c88243ddd72d3c02673945d1f

Observation 23622d0a-3561-49b8-9458-c3817c6b7f97 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.491170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.491170Z digest=sha256:206402e41e6b2b25146c5c400b89e8016ae2fb69ef125ac3c03d5bb9b57bc0b0

Observation 7d0bfe84-fb56-4288-a145-74e854cc139d · outbound

This paper cites , " * write output.state after.block = add.period write.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning , " * write output.state after.block = add.period write

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.566124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.566124Z digest=sha256:5dd423f6b94e82c14f93c7ccc611279920c9e3c608a69a090f5d57f2090beafc

Observation 2d2679c0-1349-4812-9f51-95df8bf0e25d · outbound

This paper cites write newline.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.668238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.668238Z digest=sha256:9e8a77e80e01d857b26d5f48e75d5555ead39f455475aca7a5b5aac880aff8e4

Pith citing papers

Observation e4dc91ff-51d2-4752-aa5a-47c8c4b80e89 · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:34.984840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:34.984840Z digest=sha256:40b951d236e2ce4f1954e7af951232f9f9a4706461a7d4098d0b4fc40665c2ac

Observation bf6bdab8-5b24-4a69-9e98-363b25af0fbb · inbound

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation cites this paper.

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T22:52:00.547279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:52:00.547279Z digest=sha256:001888e74231f85da26326ffff8c037cc3265a575f385ad7fa1218ea3aeb1504

Observation 1a8a7c1b-2d4e-46f7-b0c4-35065cf68ee2 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.448501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:bc7bd831a103c9872e8416fa84381915c094c8cf03622b7db767c3ab929294fc

Observation 2967d069-f60e-4f57-8581-64b9eed36b05 · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:33:02.362467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:3194a0815b59b637618229dde780400d6f6933ea9311d6ab86b592df6557063c

Observation 97ad0c2a-e9dc-4d15-8928-993a8a37b856 · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.159597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:3f491732eaadae2f9139ff129048ccb3d68396dfbf64bac4dad282715b420f2e

Observation bec402f7-7a52-459e-b962-8d50ea8d82cf · inbound

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning cites this paper.

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:03.331597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:45:21.088097Z digest=sha256:b4b4f27097818dc412605cfdda39f484ced0a448e0c3d497d79eb320a86ad4c2

Observation 36ea29d2-7175-46ed-95d2-4625e83a559c · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:10:26.498983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:2b51712dffaf35b3b1177b8b194e690bd9b4f3506dc23336945984b945bafa86

Observation ca24bd0d-f4ba-4938-9eea-6f75e42bbbf8 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:24.497209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:24.497209Z digest=sha256:e5bad75ddbb8077d13c1743c2b165061ad510131f2c043e2c39b0029631a0132

Observation 37e43759-ef2d-4700-bc61-27bd5aa2e355 · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:49.213351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:7fc237f96e8922754a9b67f0d864d703ebad68a3171d41cb427225ff97fba5b6

Observation ce49c680-a71e-42ed-919f-fd2c54d9a638 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.493453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:863c65ea7904d55ae8002f4594a9109075dcc83e2fa2d746bb98cdbc4412b1b6

Observation 39c126c9-dc8c-43cd-8c4d-2adf8647af8b · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.136829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:e0d454b9acf0918257c387b1d105c928869e0659db6ea54f5720112d342107d4

Observation 6ad6b6c6-2c8d-42bc-8c53-fde35b53ed31 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.069237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:575ef61da25e5c1b9de999911fa7633e979ec3dccac60985dd6f8c66a02468cf

Observation c51fda0a-6a8b-47ae-ae27-03d8243df2b8 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:302845471f887c135485c987528812b2048b7015b9780dae2c07a7639c5789c9

Observation bb841e23-0e1c-448e-9519-cdee7432d478 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:a90363705846a15cee9a1ee1144a02229c769867c52eaf3f8d282397916543de

Observation b77f11b1-eb43-4aaa-a7ba-0da6aa9f7156 · inbound

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering cites this paper.

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:49.554753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:57:49.554753Z digest=sha256:1bd001f444b73f86b372c0f3f4df4e70c3696b77f10f3dc98c15ca4806611f5a

Observation 07c32dab-92fe-471c-a56d-9a84f9aad946 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.203082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.203082Z digest=sha256:f25dbe25c91fe72b4c26188527efa4eb1e654e16be3eeb0d2ccda9dd2b48f167