Pith. sign in

Paper Citation Record · LEDGER

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 16 inbound Pith citation observations for arXiv:2505.22019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22019 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:24:54.668238Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:23:34.984840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.135392Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69531200-78c5-409b-9075-705dee559482 · outbound

This paper cites write newline.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:48.648661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:48.648661Z digest=sha256:f4b22afa2eb4506d131e2f986a0cbcc6172816b97316675821ad5a09b86e1a76

Observation e8ccdf16-eff3-4b06-8bc0-390b31df102e · outbound

This paper cites Qwen2.5-VL Technical Report.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:48.786792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:48.786792Z digest=sha256:23f552785ef536ea6e44659d5f3b5a74dec6b03bf0651f070a9fe09d3a5b3e8d

Observation 268315af-74ce-4a4d-8877-217f37f1bfc5 · outbound

This paper cites Benchmarking large language models in retrieval-augmented generation.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Benchmarking large language models in retrieval-augmented generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.946878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:48.906125Z digest=sha256:27a7a0af6bf90a62be305d82f628bde49d82a4544e7aaee023bfad97c1b61884

Observation 5baac46d-a7d7-4111-b649-57ccaf6cddc8 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.741210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:48.998331Z digest=sha256:be13b80ebca5316daaa91010ecf6e4d74986a49f3aab90c023e98f1e75f8b4c8

Observation cb69defb-b57e-41c1-bb83-ccf06e98986f · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.147441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.147441Z digest=sha256:c4e3d5ed4105996298d2a290fb6b5e684df31ccdcc1bc9c1416ecd2c490582e3

Observation 0af78db5-411e-4aa0-afe7-7ceefecb0d9d · outbound

This paper cites Mindsearch: Mimicking human minds elicits deep ai searcher.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Mindsearch: Mimicking human minds elicits deep ai searcher

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.352425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.352425Z digest=sha256:38d0f511c0a26602661d235780a9ff2d409c2c3ce157b4b00027accb9a1d7acf

Observation 762efeb2-6d0c-4bec-97a0-d036d833bbeb · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.490774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.490774Z digest=sha256:624ac1b11775fa68e15e71c5f53ec2cee099be75a6943397c5c663aa2974a1af

Observation e9790918-9649-4cc5-ab71-331c48752194 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.709222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.709222Z digest=sha256:e7ec9f59788c3109dede130e094fd2808d7d7015bbb02abd3bea411718a28e2d

Observation d32f2860-0616-4982-b3a8-d69e8cd7924b · outbound

This paper cites M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.927205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.927205Z digest=sha256:e39b2088fcd62d9c6bf0abcf16341f547a1cbe44eb52f7eda4faaed1642795b4

Observation b8925cca-2976-4e05-b1cd-19f0c477861c · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.065723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.065723Z digest=sha256:e4c8c8b1778bd04afa14923cc29ad2a90b979582d6c0dac1c456e517ceb634ea

Observation acb7fd87-8eb6-47a1-bd8b-4ed82d8b5d9e · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Colpali: Efficient document retrieval with vision language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.538027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:50.220936Z digest=sha256:c8c746086f89a9e9d95e154372675c7abf3bde6d64ea24730c3c541506796f5a

Observation b48724e1-3d34-4973-8f85-220a4d7e13b9 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.367366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.367366Z digest=sha256:71610ca86d82ce4c2cfeb76de0fe4f576f04d56b5d6da24d6b9cbb6002468f20

Observation 07fc34a4-9d1a-4da1-a76d-6a940277e99b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.460221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.460221Z digest=sha256:b7755cdb480836720ce6d8dd4488429f65e4397aea9bc9072dcd0fcbd8c25ddf

Observation 7ff22933-ee18-42ad-9a6e-3307aa72aa7c · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.559672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.559672Z digest=sha256:c49886694a9ec623868029d41615f17005faa30e1150d3a2ab31d3799d91a3ed

Observation 00eeeb99-9d5a-4aeb-9c23-ed590f381a76 · outbound

This paper cites OpenAI o1 System Card.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.683709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.683709Z digest=sha256:b3f8e3614c56f2e9fb3fd06f9eacc43089b684e898a25b562d5ef7ba85a09991

Observation 00e16907-51c7-4867-a3b2-3f372a8b503a · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.840704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.840704Z digest=sha256:56be0901a7d6338a48b684e72cadef16ad3549531da8798744fac2f32aebb9df

Observation f7548be0-6977-4d9b-8ed3-e540525ea30f · outbound

This paper cites DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.943846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.943846Z digest=sha256:835bc941255910b74ee6ef69c207ba7d8025bcb7aab2c42c195cf54d7ea00ece

Observation 0914aadd-44d0-486a-a3d4-d761414d506c · outbound

This paper cites Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.082163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.082163Z digest=sha256:8cbe2f703513ae9860ce6de8af0dbcf69a1fd447a4acbe91cc808cb5beeb8c1c

Observation 9ead5704-8345-4c87-852f-5f9e97d29a78 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.217474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.217474Z digest=sha256:69d00fadcca8c3d4bbcfbd349862171ff44403c319657340c6c5a4d27168ea24

Observation 8e205582-82cd-4269-8724-f25b68dd7ad2 · outbound

This paper cites Reinforcement learning: A survey.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Reinforcement learning: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.351090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:51.389300Z digest=sha256:fb35951f0e912fab080ca37f7d5b8f43804eea0658eecd0f4cbaf3836c71dd37

Observation 6d2d3ac9-3765-48a5-8255-04acfd34cbbf · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.601363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.601363Z digest=sha256:49acfe7f3a55439b20e401248d5354bb8e5e3376c20e5c96b13c17127f7040c5

Observation c2d12a3a-869d-4955-84ef-bb3ddd92a9cb · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.700223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.700223Z digest=sha256:642b4490b0edd092c7cd92741e8cb7ed305037d64f204fa028262088d25dee3e

Observation b0219256-bd67-4df6-9a69-389dde6c5c69 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.824122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.824122Z digest=sha256:a50779de05048a13d10732db5611622a6007e70303bddc0db43a0619295b14b4

Observation 11459a76-2481-48f2-b128-0ef82e496bae · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.884343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.884343Z digest=sha256:0bf4118eaf907e0cef76e2399d1d4d24ce580f39449aa7d3f4e7c33900daab8d

Observation 8b9ceca8-29b6-4506-b59b-b18646e7e1e1 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.963166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.963166Z digest=sha256:1836c5fef7ad9a6944e37e7f10c8b3634fd690c9eaf9f5998423b78c4b638336

Observation e8acbfd5-bd5a-4648-b7b9-5e14039ded2f · outbound

This paper cites Improved baselines with visual instruction tuning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Improved baselines with visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.036227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.036227Z digest=sha256:0ea20709667a31194617ebbb3f4125b974d0976b03b41d76f1f1f28bf15ec190

Observation e0f17fc1-8a65-485b-ae63-665bbe9a6844 · outbound

This paper cites LlamaIndex , 11 2022.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning LlamaIndex , 11 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.146606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.146606Z digest=sha256:303474f5f4ed9bef1655282bd8cff7a680d34b593baf8d2e6473c8832af6a08e

Observation 02dc48b8-0155-4e5a-a008-61cb42d6d12e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.248859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.248859Z digest=sha256:b719b2169013bc71e9f385fb2492ffbd619eabd0a288f30bbb08f5546de2adb6

Observation c0ed3f2c-aa1f-435a-90bb-bee0c89944d7 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.343132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.343132Z digest=sha256:8a2cb8cec7196d4bd38f8f1af9cb3b0aba56c4846b4381fe9af2e91ffd6a59f1

Observation d634e7bc-32f8-43f8-a76e-bec59d2d6849 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.421805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.421805Z digest=sha256:1d4d90811b58166eaf63fdb509be3a67cad85b0ba2c77107cefa5764ab05cb1b

Observation 2c2ea2cf-8642-469e-924b-799c5c19984d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.518896Z digest=sha256:7bee5d3132eb6ee5d837ef6b4af8cef657c4724da2ee6e0f047054277a545a9b

Observation 4c8c0d4e-b367-4663-93eb-1233a01ea3dd · outbound

This paper cites NV-Retriever: Improving text embedding models with effective hard-negative mining.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning NV-Retriever: Improving text embedding models with effective hard-negative mining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.610720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.610720Z digest=sha256:18754a7a540f3a759af3294173e35d4aa57b0b5196e89efb223ac3c6d5a7c210

Observation 4dd90c39-b902-4ffc-b08b-f7206b3b0c2e · outbound

This paper cites Hello gpt-4o.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Hello gpt-4o

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.679894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.679894Z digest=sha256:d78022bef267f6dba870fd516c1a04ca0ef9c8b96c3a1e8a0da907697ab2482c

Observation fcafe132-9e4f-4759-bc04-037d07ee81cb · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.763882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.763882Z digest=sha256:9a40b90e5b3c996c511a6643f33de58341bf489feacb47637699b3f67cf0541d

Observation d251e266-f212-41fb-bcc0-cba2aec026cf · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.861519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.861519Z digest=sha256:bcae1766bd5018e5c6d501931034e7abaa7b29d10377fa8befec543f4a54f6fd

Observation f33cf807-ebaa-485e-92ef-8f9286a003c8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.950386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.950386Z digest=sha256:6b2e30fb54cf60dc17cd5d93baeb3ae05bd87e4829f249610d3828af99b38348

Observation 65ebfb42-2e5d-49bc-9dfe-29506b85d40c · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.100544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.041498Z digest=sha256:017626a87c4121ab806bbefefb564118dfaf472e2ebd9fb984c8f15261836242

Observation d74a9cfd-9ee9-48f7-b84d-946c85862326 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.171877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.171877Z digest=sha256:a0686da13eebd72276964d7e305d781bcf1e14004d7fa3d920a2ab794d727a13

Observation b86812e1-643b-4b23-82d6-b42c127aa112 · outbound

This paper cites Reinforcement learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.809640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.264735Z digest=sha256:361b143ca7eb8108d5989e525b0cc8c7d577b392d19c16149aa7b1eadcd5e5a5

Observation ebfbe6e8-0a78-475e-9803-6fd86710633c · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Slidevqa: A dataset for document visual question answering on multiple images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.604920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.364738Z digest=sha256:2fedef766d94428f74ab3d8043f6c2cb44ff4bd982e059a96fb06c33e93e30e4

Observation 056425ae-5ace-4508-a89e-5f31bb5c7e35 · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.442043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.442043Z digest=sha256:019132d933931ff5af3a4cce552e71862d6b7c75d6e53e24bfc5ad7dae6bec10

Observation 6ebaea24-4713-4d9a-a30b-d63e2cd79fc3 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.487266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.487266Z digest=sha256:631e2b25fb0497d912b1f255d1c0ce3b599447a72c74c86478c8e12c881db33b

Observation 5bd4bdc1-02f8-4d2f-98c2-88b1d53c7328 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.586633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.586633Z digest=sha256:e31d57954d72f0888ca5095d60992d4d82c5ffa1a5a99d9175ae9fe875495f9e

Observation f15666ee-bdbc-4691-9a14-5246d2f099a3 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.713826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.713826Z digest=sha256:e74bc4707c94cdd58a49d8521304584c3418b2f3e00e091ac2ce7b1a79aaaaea

Observation e328ef64-c72b-47c8-8eda-6451b5932802 · outbound

This paper cites Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.810076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.810076Z digest=sha256:9be09573914ab1057f5ec75085f5ad74640e46306b304049323beb526fa4b677

Observation 04b2de5d-2172-4ee4-baca-17c71b3a49fa · outbound

This paper cites Rule: Reliable multimodal rag for factuality in medical vision language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Rule: Reliable multimodal rag for factuality in medical vision language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.380239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.914447Z digest=sha256:2b707a8f66769ce05cf0625dcd56f894af84560b17979bdd00649861a67732b8

Observation 3ec533b4-47bb-425c-bc96-4f0808b4965f · outbound

This paper cites Qwen2.5 Technical Report.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.021458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.021458Z digest=sha256:bb3522fdf5e671132f9fb53b14c76997c6bd1dd8a65f63fae2cfbc196782deff

Observation 161fa363-a68a-42c3-8582-5de71460e329 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning React: Synergizing reasoning and acting in language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.132125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.132125Z digest=sha256:be659a485016465caa7f483382eb140c12c0c3f0d4cf8ab4c5524245abc52969

Observation 512576cf-8bdd-4828-88db-3d113f7ce37f · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.235680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.235680Z digest=sha256:9f9cde720bb305ce1163cc45f839ab25559f5da215251f852cb178ee424db240

Observation 59e799c9-745e-4f80-8775-79913f385c70 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.345474Z digest=sha256:c95d331f907d3d0e0eac0da54acb01bbce9d31cdbd5199b0cf08e542ed988982

Observation bdcf5e24-caf3-41f1-8888-177af150b37d · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.427481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.427481Z digest=sha256:6f1fb06998a18e28da979f3c290c1cd448229fea07f56606b40caf15039478b9

Observation 23622d0a-3561-49b8-9458-c3817c6b7f97 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.491170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.491170Z digest=sha256:ad54c221f5ac4e0cc23dbe3065ed843d1dabeed6d3b0da3d0d022a17d9cf8444

Observation 7d0bfe84-fb56-4288-a145-74e854cc139d · outbound

This paper cites , " * write output.state after.block = add.period write.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning , " * write output.state after.block = add.period write

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.566124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.566124Z digest=sha256:6306e47d1ea4bda961ad3b53de8a23aa2ee8d228551f32bcf1670772105017a2

Observation 2d2679c0-1349-4812-9f51-95df8bf0e25d · outbound

This paper cites write newline.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.668238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.668238Z digest=sha256:a60b5c52d817c8ea5a9479446b729de663c4069ff737da5eb1119299f978b7ce

Pith citing papers

Observation e4dc91ff-51d2-4752-aa5a-47c8c4b80e89 · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:34.984840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:34.984840Z digest=sha256:f06f71086bceb321eb805684468b176927515251ee4dba0cefad3109a997d302

Observation bf6bdab8-5b24-4a69-9e98-363b25af0fbb · inbound

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation cites this paper.

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T22:52:00.547279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:52:00.547279Z digest=sha256:6fbd8ceb5676b91936a7260f147e34432bd6721732958fdd7c33c41b695bdc0e

Observation 1a8a7c1b-2d4e-46f7-b0c4-35065cf68ee2 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.448501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:f853f6e656bda265be208bfe72bbdc7d88ce1f0c4588cbea5f5eeecf017e6609

Observation 2967d069-f60e-4f57-8581-64b9eed36b05 · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:33:02.362467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:1da840bddda510bee97f0fd4571aacba25a09974a8ea872834daf8810935b828

Observation 97ad0c2a-e9dc-4d15-8928-993a8a37b856 · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.159597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:e1ea796f0c04c17e613bc2d5f6484c261dcfe310f8cbc808b9acfe6cf303c02e

Observation bec402f7-7a52-459e-b962-8d50ea8d82cf · inbound

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning cites this paper.

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:03.331597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:45:21.088097Z digest=sha256:e35dea9bc91396a7a1ae59434f3a74af6c8dc8a9e9c94695110f753841c94287

Observation 36ea29d2-7175-46ed-95d2-4625e83a559c · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:10:26.498983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:ffe03b1ed8caadd8d92cadf560415c5db322744894017b97adfa4c9a40dfcd05

Observation ca24bd0d-f4ba-4938-9eea-6f75e42bbbf8 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:24.497209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:24.497209Z digest=sha256:2b77024ac6b62d993691e80a3dd65d84c8e16a825266d5009566265ce5b1f937

Observation 37e43759-ef2d-4700-bc61-27bd5aa2e355 · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:49.213351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:c52d2162bc46c41e35efe4950848f5ff9628d8813a2982a4bfdb61d614101a74

Observation ce49c680-a71e-42ed-919f-fd2c54d9a638 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.493453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:513fdeb30b07e05672b395a4ed29132e98b6bb3d6083f757dfcba5d6a1baf297

Observation 39c126c9-dc8c-43cd-8c4d-2adf8647af8b · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.136829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:c0d4803a183669749acad9b148b30ea4b602ede7c99e7babdcd24ec3b346f52d

Observation 6ad6b6c6-2c8d-42bc-8c53-fde35b53ed31 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.069237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:c8fb38406b83d6c8cc0363f0952006bc1e30ddfb508c22e549d4ef59e2461390

Observation c51fda0a-6a8b-47ae-ae27-03d8243df2b8 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:0e36b5f94d3175697ec703ed98e9f083e436007c653c20017e5e6d5e2454fe0f

Observation bb841e23-0e1c-448e-9519-cdee7432d478 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:1f00a6123a56310d654599d7eb91024a63573e2e59d61d686aff7c46338d6627

Observation b77f11b1-eb43-4aaa-a7ba-0da6aa9f7156 · inbound

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering cites this paper.

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:49.554753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:57:49.554753Z digest=sha256:1008af35f013f94b9b3627ddbf7e6ad55987687e33476696b161f848f6c18a46

Observation 07c32dab-92fe-471c-a56d-9a84f9aad946 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.203082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.203082Z digest=sha256:db2bbded883a291d285bc106058cb0736e952a5810b7d2c55fc07dc273252538