Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-15T14:43:21.866055Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 4 inbound Pith citation observations for arXiv:2603.05256.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-15T14:43:21.866055Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:19:52.487643Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T13:15:50.537044Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0af5cc2e-73f6-434d-b2b5-9b9eb0a27ade · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd4f9ae-c567-4359-a94a-2355379060a8 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9276f29f-33bd-410b-af3e-8da03a405f2f · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30d93720-acf0-4999-9137-0f9bc717ffea · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff88c41b-3399-4216-8557-de964ce69775 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Automated Curriculum Learning for Neural Networks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f15c67-efb6-4c83-8562-bc0caaa8c014 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum KAT: A Knowledge Augmented Transformer for Vision-and-Language
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ae7642-29a6-439c-b041-bdd165e1a10f · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Belongie, and Oisin Mac Aodha
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa81d54c-c886-457a-b8a6-020ae65de71a · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Open-domain visual entity recognition: Towards recogniz- ing millions of wikipedia entities.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd21357-4079-4621-96bd-7f946cac61ec · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Ross, and Alireza Fathi
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d1301b-0bb4-41c2-9127-2b4cf3e91d3f · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unsupervised dense information retrieval with contrastive learn- ing.Trans
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4d6b1b-9181-4a1b-968b-c05b004996ec · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5398e1cd-7782-4cc8-8d7c-d5d88e78bc37 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98529806-ac7b-4936-ac38-3b33847399ab · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4504f9-1ac3-498d-a5ef-d9ff8b9a0876 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Jabri, Trevor Darrell, and Pulkit Agrawal
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a35a32-a5df-4179-96f1-e87b8be39445 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Retrieval Augmented Visual Question Answering with Outside Knowledge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e15ba6-34e2-4b2b-899b-1750464a0388 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Improved Baselines with Visual Instruction Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d3cd3c-64cc-4712-a6d5-4ed13e5abdab · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Ok-vqa: A visual question answering benchmark requiring external knowledge.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c699fd-0156-404a-aebd-9023a4f53c2a · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1b24c4-8c11-4e82-916b-4f9b53c66b66 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a5215c-499d-4598-a426-66dd4594c653 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Teacher algorithms for curriculum learning of Deep RL in continuously parameterized environments
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8e8219-715a-466d-bd38-2561b2015aad · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum RoRA-VLM: Robust Retrieval-Augmented Vision Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e539cf-d27a-478c-a954-38630ba99d89 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e289fa-11bc-4a9f-8f27-182287debe3d · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede4a04b-f298-4062-b4ce-3454ece8599b · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Da-dpo: Cost-efficient difficulty-aware preference optimization for reducing mllm hallucinations.arXiv preprint arXiv:2601.00623,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584dec0f-5852-4129-a3ae-6ebc62f2ab84 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e534df68-a116-4a92-8f05-d92c6375fc09 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c703e467-541e-4475-8e23-6bbed5b26eeb · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8712bc-ab4d-4959-8816-63fc7a3e0b9a · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Denny Vrandeˇci´c and Markus Kr ¨otzsch
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af17a9cf-6d77-4038-a236-4966c15d81a9 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum URLhttp://dx.doi
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a6461c-7480-40a2-b2fc-edf7a0cd197b · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f496a94f-5f40-4ad5-a35d-e4b4a4238497 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Zhang, Zheren Fu, and Zhendong Mao
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0fb5d2a-2a82-4e36-9101-604b24e143e7 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum de Ara´ujo, Bingyi Cao, and Jack Sim
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c7f8a3-dc2b-4c08-8da0-8a6d1fae7926 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum MMSearch-R1: Incentivizing LMMs to Search
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1d35240-cf00-4363-9214-d66c64820bb5 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum A Simple Baseline for Knowledge-Based Visual Question Answering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909a70fa-6c91-4148-aec9-d57a7798f601 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EchoSight: Advancing Visual-Language Models with Wiki Knowledge
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a602055-d793-4878-b3ac-157051c6e352 · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6797923-e4a0-476d-9806-a043a665a98b · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 147f527f-1a30-4f3f-b68a-29f9c48c84ad · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum A curriculum learning approach to reinforcement learning: Leveraging rag for multimodal question answering.ArXiv, abs/2508.10337,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 687f5520-d56f-4d6c-b7f4-50f7f249cdfe · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum The CLIP I-I is the retrieval with the visual similarity score from EVQA-CLIP 8B only
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d86120-d739-439f-84b7-a9b979c2faae · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum As shown in Table 7, our method requires substantially fewer training samples while achieving superior performance
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a51a36c9-eba9-480d-9a97-f563dcf4e12a · outbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4c332e-dab1-4a1a-80b7-36f3ab08e987 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2ec3ef4-9853-46aa-85f3-7dc53a3d6866 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 693e040e-3774-403a-83e3-8afd4206c81f · inbound
UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d15d2f-478c-4f03-942e-75e496b055cc · inbound
UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.