Pith. sign in

Paper Citation Record · LEDGER

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 4 inbound Pith citation observations for arXiv:2603.05256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.05256 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-15T14:43:21.866055Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:19:52.487643Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T13:15:50.537044Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0af5cc2e-73f6-434d-b2b5-9b9eb0a27ade · outbound

This paper cites Qwen2.5-VL Technical Report.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:3951dacb58d853636c52c71f1e7bf4e29a1cdf8312564f94e2c87fbe8eec12b6

Observation 4fd4f9ae-c567-4359-a94a-2355379060a8 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:ffb9d5cd0918e5fe2ae98816a2892655f7ddf57f423800357c15a02940ded648

Observation 9276f29f-33bd-410b-af3e-8da03a405f2f · outbound

This paper cites an unresolved cited work.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:9dd18a19060218889737f69ae6656e39dad52f8fd6dc74f6d22666a188c313c2

Observation 30d93720-acf0-4999-9137-0f9bc717ffea · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:a851e7856b143e9c941f9a2ab703742041501a750d2cc34856b51e2ec65f8554

Observation ff88c41b-3399-4216-8557-de964ce69775 · outbound

This paper cites Automated Curriculum Learning for Neural Networks.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Automated Curriculum Learning for Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:185972e7d66b110888f467a2eda70c11e04fda6388d3303a2038147bfb3e63c5

Observation 30f15c67-efb6-4c83-8562-bc0caaa8c014 · outbound

This paper cites KAT: A Knowledge Augmented Transformer for Vision-and-Language.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:baac5713b397bc2781e74a538efd93827c39f05d64e3156a149675d9fa1633e0

Observation 45ae7642-29a6-439c-b041-bdd165e1a10f · outbound

This paper cites Belongie, and Oisin Mac Aodha.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Belongie, and Oisin Mac Aodha

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:d5516ef772318259e5ce3eeba1fb9135caa328189d27a21921d921029e380613

Observation fa81d54c-c886-457a-b8a6-020ae65de71a · outbound

This paper cites Open-domain visual entity recognition: Towards recogniz- ing millions of wikipedia entities.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Open-domain visual entity recognition: Towards recogniz- ing millions of wikipedia entities.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:99d6914b2572668bc0b657469543d05597dcae801b36d3c297be28189cf5de39

Observation 7fd21357-4079-4621-96bd-7f946cac61ec · outbound

This paper cites Ross, and Alireza Fathi.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Ross, and Alireza Fathi

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:f3d5b52d35573ed51da0cc917525cbbc8af8ef7f5192f039425c15dc5976f873

Observation 15d1301b-0bb4-41c2-9127-2b4cf3e91d3f · outbound

This paper cites Unsupervised dense information retrieval with contrastive learn- ing.Trans.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unsupervised dense information retrieval with contrastive learn- ing.Trans

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:2427e4b01d8e90b6e877f721d92dc00f9ddefe0011dbe19a5b34447bb21eaf53

Observation 4a4d6b1b-9181-4a1b-968b-c05b004996ec · outbound

This paper cites Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented Generation.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:717a8acef2c9d8f7424cb292af19ff3530f04b954bb7ec240a1e2d545ce4d4d4

Observation 5398e1cd-7782-4cc8-8d7c-d5d88e78bc37 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:5be795112dd3276e6cfd456d018be73d9b092573cd75cac0386b5a329b71aec1

Observation 98529806-ac7b-4936-ac38-3b33847399ab · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:78776a58d24fab3ecba49b9fca99b42b526253883888ee6ce9db3018735ee7d2

Observation 0a4504f9-1ac3-498d-a5ef-d9ff8b9a0876 · outbound

This paper cites Jabri, Trevor Darrell, and Pulkit Agrawal.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Jabri, Trevor Darrell, and Pulkit Agrawal

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:db60f0bd832342d1d40e2bcd2a55f3258bbcf25c926e377f5670921ae09b8425

Observation 14a35a32-a5df-4179-96f1-e87b8be39445 · outbound

This paper cites Retrieval Augmented Visual Question Answering with Outside Knowledge.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Retrieval Augmented Visual Question Answering with Outside Knowledge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:e73a89a5b628a6e0e2f0f3a1d60237f4729cd8045ea9e012edc53791fc8db299

Observation a2e15ba6-34e2-4b2b-899b-1750464a0388 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Improved Baselines with Visual Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:633dc93c9981295dd42c8b9bcc14b2bbb1f7e99021d45a0e879c758c1a61ed2a

Observation 53d3cd3c-64cc-4712-a6d5-4ed13e5abdab · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Ok-vqa: A visual question answering benchmark requiring external knowledge.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:cddf3d6f7ce6a5c0e3fcb66afd4fa896e0abd91e81ad67f8f3f4f8209fc071f5

Observation 43c699fd-0156-404a-aebd-9023a4f53c2a · outbound

This paper cites an unresolved cited work.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:2712e694ef806fff0a94907f16b0f20cfcd7131304980c3dcb97eaeafba1d3c8

Observation df1b24c4-8c11-4e82-916b-4f9b53c66b66 · outbound

This paper cites Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:2f4ec23cf7a109991e5b69483adbbc793ca1ce79d4f8d39fdcdccc1a4b016750

Observation c6a5215c-499d-4598-a426-66dd4594c653 · outbound

This paper cites Teacher algorithms for curriculum learning of Deep RL in continuously parameterized environments.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Teacher algorithms for curriculum learning of Deep RL in continuously parameterized environments

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:409ac1ccee6310acc2e862acdfde9c5a7315bbb44e4a7528b84885f922514833

Observation 1b8e8219-715a-466d-bd38-2561b2015aad · outbound

This paper cites RoRA-VLM: Robust Retrieval-Augmented Vision Language Models.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum RoRA-VLM: Robust Retrieval-Augmented Vision Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:a706ec1bf35929e20c326d1e139419108752cc362e33023d6bf05add8a39edd9

Observation 42e539cf-d27a-478c-a954-38630ba99d89 · outbound

This paper cites Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:c44e258266ea70b4f0719a5c67c1fac892040241b4df3ff7f87ad34c2d82bb3b

Observation 25e289fa-11bc-4a9f-8f27-182287debe3d · outbound

This paper cites NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:66600373584eed11c161682e954f38a4394dc36afea026d8a4e4389de6c47182

Observation ede4a04b-f298-4062-b4ce-3454ece8599b · outbound

This paper cites Da-dpo: Cost-efficient difficulty-aware preference optimization for reducing mllm hallucinations.arXiv preprint arXiv:2601.00623,.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Da-dpo: Cost-efficient difficulty-aware preference optimization for reducing mllm hallucinations.arXiv preprint arXiv:2601.00623,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:da1cd60c3f6516055991edd6a30ec2a4f2c4c5724e9fed6e8a0288f5eba6ecb6

Observation 584dec0f-5852-4129-a3ae-6ebc62f2ab84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:ec974574d11fe77aaffda3888da0b94e9683f27a2dce826cc85527afba34914a

Observation e534df68-a116-4a92-8f05-d92c6375fc09 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:d272f1456711540494d38a888808c86d729e38dfee7496de067cd28606caf3df

Observation c703e467-541e-4475-8e23-6bbed5b26eeb · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:f251f4b4682978180f2b60223ef43f9870046cd08fa00d906551ce1f0041cbee

Observation 1b8712bc-ab4d-4959-8816-63fc7a3e0b9a · outbound

This paper cites Denny Vrandeˇci´c and Markus Kr ¨otzsch.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Denny Vrandeˇci´c and Markus Kr ¨otzsch

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:4af04d569380a3458caf4bc1ac130c7e1a452b0633adb7fd568eb2c1180dbb47

Observation af17a9cf-6d77-4038-a236-4966c15d81a9 · outbound

This paper cites URLhttp://dx.doi.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum URLhttp://dx.doi

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:3e686b653814ac468f460daad207d3f8640b52f1b470e358b16eba9f3944045c

Observation 14a6461c-7480-40a2-b2fc-edf7a0cd197b · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:67904d8f9e9c6b3daea6ca79732a026ece6b90c8f296e5dcc13b6e52017ec74d

Observation f496a94f-5f40-4ad5-a35d-e4b4a4238497 · outbound

This paper cites Zhang, Zheren Fu, and Zhendong Mao.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Zhang, Zheren Fu, and Zhendong Mao

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:77c4d6d51e194fe4377441471aed44995fc88c7990686beab124c65d9ab3ef18

Observation d0fb5d2a-2a82-4e36-9101-604b24e143e7 · outbound

This paper cites de Ara´ujo, Bingyi Cao, and Jack Sim.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum de Ara´ujo, Bingyi Cao, and Jack Sim

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:2dd262c0471223e24e6a3a51b523e3ce49cfcc0e7f13c67b912bcd682cf4a58a

Observation 09c7f8a3-dc2b-4c08-8da0-8a6d1fae7926 · outbound

This paper cites MMSearch-R1: Incentivizing LMMs to Search.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum MMSearch-R1: Incentivizing LMMs to Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:484b2c94a061ba21f28d47c512983d587f20c55ef95cffbd148ddedbef3f7710

Observation a1d35240-cf00-4363-9214-d66c64820bb5 · outbound

This paper cites A Simple Baseline for Knowledge-Based Visual Question Answering.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum A Simple Baseline for Knowledge-Based Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:248cbe6a617c677b3ba6e94fd1478229af2c65e3b874768fe53b132a088518b7

Observation 909a70fa-6c91-4148-aec9-d57a7798f601 · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:dc8d7325679d0eb10cbef1580d2dde18bb62e593920cda298356eadb1e15ca2e

Observation 1a602055-d793-4878-b3ac-157051c6e352 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:5d46bbce1951042cd8f0b3de47acfe7940be97cea078cf176a0589278c69bf1c

Observation d6797923-e4a0-476d-9806-a043a665a98b · outbound

This paper cites VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:0be3594d2e724fdedc92f3bc94d75dc60197294785f4634b9f1a74c9ad22454b

Observation 147f527f-1a30-4f3f-b68a-29f9c48c84ad · outbound

This paper cites A curriculum learning approach to reinforcement learning: Leveraging rag for multimodal question answering.ArXiv, abs/2508.10337,.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum A curriculum learning approach to reinforcement learning: Leveraging rag for multimodal question answering.ArXiv, abs/2508.10337,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:49025029ef609d5c0d787f2b46b1f9346b3a6f990a09ff1a107cb96b1537a9fd

Observation 687f5520-d56f-4d6c-b7f4-50f7f249cdfe · outbound

This paper cites The CLIP I-I is the retrieval with the visual similarity score from EVQA-CLIP 8B only.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum The CLIP I-I is the retrieval with the visual similarity score from EVQA-CLIP 8B only

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:93b060e63f98e06592d0e57180c3464aa19757c9bb60748a7dd0f40aa365e059

Observation 42d86120-d739-439f-84b7-a9b979c2faae · outbound

This paper cites As shown in Table 7, our method requires substantially fewer training samples while achieving superior performance.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum As shown in Table 7, our method requires substantially fewer training samples while achieving superior performance

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:8e7539312e875d22ec054df11365b71555679e357aeca18eb21f51529f0336d7

Observation a51a36c9-eba9-480d-9a97-f563dcf4e12a · outbound

This paper cites an unresolved cited work.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:99d3f30a1c40d0203ce3ea880f6f568019d871ce62045f9eb20c5320fd3ea158

Pith citing papers

Observation 9f4c332e-dab1-4a1a-80b7-36f3ab08e987 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:35.996569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:42a36b14ad59e924431b177a60b6a7a8c1d570d69d33db5f133f39ac0c8089cd

Observation d2ec3ef4-9853-46aa-85f3-7dc53a3d6866 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:9994fa43898df2f20096693fdc208ca3fb12ecf5257bb8ce6534fde413db8d03

Observation 693e040e-3774-403a-83e3-8afd4206c81f · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:16.714259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:16.714259Z digest=sha256:16171e784342d8ef064dfd01c4dba4bf328afa33b94161fe54ba59fd72aaa3b4

Observation a8d15d2f-478c-4f03-942e-75e496b055cc · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:52.487643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:52.487643Z digest=sha256:2f3bb28a6c24031ebf0f9b5095819a2905b89697ce4749d04871162a553358de