Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:05.282142Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 8 inbound Pith citation observations for arXiv:2506.07905.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:05.282142Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:07.475851Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:39:46.104997Z
100 of 125 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a2f85b3b-0818-41ce-9958-2cae1fc7f270 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978c1cab-a1dd-4444-89c9-6ad7b5172600 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Openai o3 and o4-mini system card, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4843fe97-3092-45f2-bf8c-925fce292f34 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9fbe446-e9d6-4e54-8d3b-2e4157656162 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03394a63-8b2d-453b-831a-db6468728ed7 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb95a9b-1dc8-4365-8085-9f50c9b4469e · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning s1: Simple test-time scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822dc9e1-c613-4d34-ac15-2d56b77e2daa · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495b0a75-d03b-463d-83cb-f010a3d40fbf · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5894646a-2c5d-4728-b863-800dbed0db4a · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32720cc-9242-4df5-9434-3ef7c2fcbb91 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65f2783-cb81-4d0b-81c6-d4f0b5b9c89e · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f93ccf-1a8f-4b7b-a910-5f065c564f82 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853b1ea0-8045-4c13-85fe-1bc958a89ac2 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9527092-abfe-40c2-90fc-066cb510a9a0 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be5a1d3-a000-4c5f-9957-538d64e23047 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Noisyrollout: Reinforcing visual reasoning with data augmentation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931c915c-c4eb-46fc-a16f-3c8a5d0e56df · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4360f939-c564-4e2d-b0db-dd4d8b278137 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734dfa59-790b-4812-9760-d728fd635084 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43752052-9380-4f73-bef0-41ec3dc11c36 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VideoChat: Chat-Centric Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad8d57d7-e276-4a9d-acab-25b25134dd35 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b78e31-cdda-46c9-8643-9501c5f25b5c · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bb33cbc-892d-4c18-906c-d4aadfe706ef · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3691229f-741c-4818-8bb2-a29dd3b1b139 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Cogvlm: Visual expert for pretrained language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca83f3e-479e-4246-af88-9e0ac8d0177e · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Sharegpt4v: Improving large multi-modal models with better captions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773ba099-657e-40ee-b8de-4024deea19f1 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2523b7a-b2f7-45dc-9ef8-6112f043af83 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea755578-ed6e-48c1-a8fc-4b270399c075 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved baselines with visual instruction tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e4bad6-63ba-4c71-bc04-1cd724b58701 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85856486-bfae-44a9-a157-c0638ef44a41 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 866e57e5-febe-4c62-a48f-569875f5fdc9 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d41cacd-b65d-4fad-ad0e-98d679aad327 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743ef82d-e42c-42e6-9bdd-0f9672ef23e6 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f02f92-f888-404a-bb3a-02e0f1357460 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a55093e7-af8c-4648-80c8-294a51b3d510 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning GPT-4o System Card
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f62a4c-2475-4afe-8820-7f6546cb2512 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Claude.https://www.anthropic.com/
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f81c7e-41dc-4c9f-a147-ca620750ba25 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Grok.https://x.ai/
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580d4b2b-7a81-4cf8-a2d8-cf1b67f926ee · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217880b1-42c6-488a-b0f8-96e668dc1cee · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373b2e0a-d748-48a8-967e-69b399de2ec1 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Multimodal Chain-of-Thought Reasoning in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00f1f6c-04fc-4e3f-8008-a8d0088b3fbf · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed55363-e79c-4b63-834e-d9e1eb30dab3 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Layoutllm: Layout instruction tuning with large language models for document understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b587f56-f0a4-4f1f-a570-c273cdc5878a · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5793574b-1f18-435d-8eb9-a6f95939c91f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e96661-8911-472f-b823-13befa0f4b47 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Compositional chain-of- thought prompting for large multimodal models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4093cbf7-d757-4e22-86d6-75ffac09711c · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88c960c-7624-418f-870a-1bde9f3e909d · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 722c6d23-e56b-4f1e-8965-a73b0ea02771 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual chain-of-thought prompting for knowledge-based visual reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b2136dc-12f3-45aa-8bda-ac0ed1d08a1c · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b21cc11a-7cb7-4090-b1d7-b43f9e8a02c6 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38fca845-86fc-4356-abf4-22a29e52629f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e8ca4e-e792-4925-8635-57e9b487c9aa · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d226f46e-6381-44ed-a5dd-84a08c35a2d9 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f129975a-ae54-4a41-a608-82aca32d2b88 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.arXiv preprint arXiv:2504.14642, 2025
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3c6233-1240-410d-b735-770dffa707f5 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Compile Scene Graphs with Reinforcement Learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0606e7-931c-4f6b-a038-f2783832d086 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900592f3-cb75-485e-8f6d-57cc9329113b · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4216eaa-f78b-4267-8cf5-86dd6946c4ec · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 103c64de-bd83-4be8-9b3f-cd361c7ee95f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d120e7ae-903b-489a-aa89-14f38d5d8a2f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Perception-R1: Pioneering Perception Policy with Reinforcement Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37fd741a-0b79-46df-9e00-a1c67bf22259 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7679a474-d1ab-4409-abd9-9c4aaa7000b6 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6348801c-8c22-4cb9-8943-fd8dbb1ab974 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bc375c-4e64-45da-a171-f57595902bb2 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d157d2d5-3671-4249-8e5b-5e86a6654f91 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.arXiv preprint arXiv:2504.03724, 2025
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4217f07e-78ef-41aa-84e8-8aa534724142 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc15cf8-7aa9-4a9e-8c30-3b2432984ada · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c581d90e-8ee5-493f-9b53-e0e95a7b9f11 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Microsoft coco: Common objects in context
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d04fc18-ed40-4c4a-a6e7-95235979c46f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Segment anything
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89861dc0-b9bf-4763-a63b-48c10926ab6c · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e381a6a3-9853-40d7-8bf7-9d101078f2ce · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3867a4-b03c-4462-b451-93b616482945 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Dual-glance model for deciphering social relationships
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca7b2491-d32b-4813-98bb-6b7c5dc20abb · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Towards vqa models that can read
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291d4e8a-b359-415c-b459-b6c1c0a85b02 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Docvqa: A dataset for vqa on document images
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a10dcfd0-96a0-4c4e-8c9c-f3242c3acade · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ocr-vqa: Visual question answering by reading text in images
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 068b58b0-13bb-4fbd-9d9e-84d6034677ae · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9aa5595-ee68-44df-bb08-01cfb3932d8d · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d62ce74-5137-4911-8044-048da7265ebd · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b06efcd8-19f4-4144-843f-9821c78a36cc · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning A diagram is worth a dozen images
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabeb259-7914-4223-8ee5-9a69dae88dbc · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cca9462-d10f-4096-afbf-ed6e80f2ff0e · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de39bb08-53f5-4639-9513-e13008944536 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Quality at a glance: An audit of web-crawled multilingual datasets.Transactions of the Association for Computational Linguistics, 10:50–72, 2022
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 453ec9e4-fd2a-452c-b484-b24161f7711f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2217614f-6fb3-46a2-bb53-017653874857 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31eea80b-7b86-4b09-a24c-f21b2b5d36fa · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Measuring multimodal mathematical reasoning with math-vision dataset
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc998b26-193b-44a7-8a6f-d61669a8b7a9 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5bc9403-a0d6-4efd-b92f-bc0e7a252ec1 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d11d3584-06d5-4374-867b-ff60a6a0b96f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2583034b-a82f-497d-8e02-8cfb62091901 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6b901e-2631-4aa3-8984-2716b6426aa0 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6017241e-0ff7-461c-8c6f-b5ba0517dc04 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3dbcfa2-99e2-44f3-b95c-5f8a0ae48840 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98464604-acdf-499c-acf3-4e38f8e7e286 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24ffa82-7d1f-4a5a-818f-553dd3864afa · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d2368b-9ee9-4596-bafa-2c10545d8f6f · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed61cc2-61ea-49e4-bd12-eddbad3687f9 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c09f3aa-a7bc-4afb-93a0-874cd4d96168 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0738cffd-d5a9-412b-a106-c939684827a1 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeek-V3 Technical Report
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc8fc6f-9d29-414e-90d5-719c3458ef99 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Easyr1: An efficient, scalable, multi-modality rl training framework
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40909008-8a02-4e26-98bd-9a1ed8bbe2eb · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 522bca22-8ced-4f8f-b805-8fd38471bdb4 · outbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Unresolved cited work
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db245629-0bcf-49c4-9b7c-5aa648952409 · inbound
Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f9e82d-4567-4976-9de0-b0ba9a4a3bb7 · inbound
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e25eed-7b27-47cf-834e-9b44876c6c93 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 238
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0664b863-da57-4521-b0a4-9e6706ad160d · inbound
$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b81be2-5b5c-4e7c-a979-e9098ce6c857 · inbound
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6983eb48-fcc7-4ed4-9e77-1de6f3922bfd · inbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09b8ffd8-cfb9-4af7-b9d0-3ed26b71c759 · inbound
VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a38f450c-7479-419e-a93a-2d14d92ed709 · inbound
Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.