Pith. sign in

Paper Citation Record · LEDGER

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 8 inbound Pith citation observations for arXiv:2506.07905.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07905 v1

Coverage vector

measured 100 of 125 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:05.282142Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:07.475851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:46.104997Z

Reference resolution

100 of 125 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2f85b3b-0818-41ce-9958-2cae1fc7f270 · outbound

This paper cites Qwen2.5-VL Technical Report.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.910573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.910573Z digest=sha256:2e3a6cffa33be8f0448d16cf71214070c55a95cc239c814003e2e2dd0b72d525

Observation 978c1cab-a1dd-4444-89c9-6ad7b5172600 · outbound

This paper cites Openai o3 and o4-mini system card, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Openai o3 and o4-mini system card, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.914834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.914834Z digest=sha256:a0b9cabc6f914661e81b6a19664a322c3adb4063e27e05b87d1f3f9097413e47

Observation 4843fe97-3092-45f2-bf8c-925fce292f34 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.918664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.918664Z digest=sha256:24acc7e8f4b5db3e6d88b456df23dc400c47c8d39ee9c49c9be59204f084a326

Observation f9fbe446-e9d6-4e54-8d3b-2e4157656162 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.922660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.922660Z digest=sha256:02a3604c2dda79f7fe3985ff315c7d8cbeeddd1289f24367cc5d47c6bf99ad84

Observation 03394a63-8b2d-453b-831a-db6468728ed7 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.926476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.926476Z digest=sha256:2bdf9fbda0c17263b28c82a77a0c83c35ba8a4eea205e661bcc9c249154d6c89

Observation 3fb95a9b-1dc8-4365-8085-9f50c9b4469e · outbound

This paper cites s1: Simple test-time scaling.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning s1: Simple test-time scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.930282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.930282Z digest=sha256:1a30a5f1d747c75fc90fd7d823b1bc746a2ff99670c11dae96c9f51df4305b3d

Observation 822dc9e1-c613-4d34-ac15-2d56b77e2daa · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.934233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.934233Z digest=sha256:f10b41cc635c3359ae6cc0f17ec4cdefddc56b599c9ca1e0b07e3dac217b481b

Observation 495b0a75-d03b-463d-83cb-f010a3d40fbf · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.938064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.938064Z digest=sha256:20bdbca25c625fe149802cfd1c9d6ae880dba9331c9058971cdbd296c26df145

Observation 5894646a-2c5d-4728-b863-800dbed0db4a · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.941900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.941900Z digest=sha256:6a59b2643defa4186372b7f5f13602919dc4829675f48c1107c4141a931c0fc3

Observation b32720cc-9242-4df5-9434-3ef7c2fcbb91 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.945637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.945637Z digest=sha256:690b16fa7bb661a9d998b189f14d51d157041b463b347d47591dff273e9cbb48

Observation b65f2783-cb81-4d0b-81c6-d4f0b5b9c89e · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.949338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.949338Z digest=sha256:c2a7bdbddb121295f834166be201e3859a82f413ae6b43901eea8f2ece275520

Observation 52f93ccf-1a8f-4b7b-a910-5f065c564f82 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.953540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.953540Z digest=sha256:07f55c97744321f8dabad3e70d8428e97427396d93b629e05f06704bb490498b

Observation 853b1ea0-8045-4c13-85fe-1bc958a89ac2 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.957024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.957024Z digest=sha256:c33a6cc339f15c877113c533c381a3cef68092071c92f89379485ea849a8a44b

Observation c9527092-abfe-40c2-90fc-066cb510a9a0 · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.960499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.960499Z digest=sha256:8a269f05c04989cbbfda22cfdb3f6b355f880edeaa9eef7ed83ef3b099778e8b

Observation 2be5a1d3-a000-4c5f-9957-538d64e23047 · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.964036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.964036Z digest=sha256:3cb703b31c361ae4b220dbe6e927a350c3c31607753ab26a0c9c4857243a5e2c

Observation 931c915c-c4eb-46fc-a16f-3c8a5d0e56df · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.967572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.967572Z digest=sha256:49513213ae249a4795b28523e28f0676f71ddf387c78e05b990e872a2e08cffd

Observation 4360f939-c564-4e2d-b0db-dd4d8b278137 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.971025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.971025Z digest=sha256:333913f8e03894e8725b0ecc82c63c71b52ac85a586b58f0e072b5bf2ff3c5fe

Observation 734dfa59-790b-4812-9760-d728fd635084 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.974789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.974789Z digest=sha256:85f47e2878dc9d8560b2bea245318356a126439ad1b6ed62c99a37f5a92e8769

Observation 43752052-9380-4f73-bef0-41ec3dc11c36 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.979169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.979169Z digest=sha256:aad19488a8e8b4debbcd57899c8509befc648f010729438161da69cb5ba4e8a2

Observation ad8d57d7-e276-4a9d-acab-25b25134dd35 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.982714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.982714Z digest=sha256:734594be816c1c06f0016a478a56a0411d9dbc25d0e25d29855a7cf16dd21405

Observation 53b78e31-cdda-46c9-8643-9501c5f25b5c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.986725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.986725Z digest=sha256:6a9aee75a3af6f57294de8a006b55b6fe8b9e15c7236ba7eb907b483ea2f2aa9

Observation 0bb33cbc-892d-4c18-906c-d4aadfe706ef · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.990004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.990004Z digest=sha256:dfd3df6baeefb4f5b59545940f8fb13addc3b110d4f6d3108668a095031164bb

Observation 3691229f-741c-4818-8bb2-a29dd3b1b139 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Cogvlm: Visual expert for pretrained language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.993533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.993533Z digest=sha256:39f1e0e9930ad0412446f11d57de34b0d2aa96598c59c08362ed7d4806614ec5

Observation cca83f3e-479e-4246-af88-9e0ac8d0177e · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Sharegpt4v: Improving large multi-modal models with better captions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.997106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.997106Z digest=sha256:af7cc5b83a5dd9e66b9cb18a5e1a2f4121a109b3d44c12861d554e801539f457

Observation 773ba099-657e-40ee-b8de-4024deea19f1 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.000341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.000341Z digest=sha256:e7464c7ea36a21485ee311b161d39939abce9420f9ec98fad29cae27c22fbca8

Observation f2523b7a-b2f7-45dc-9ef8-6112f043af83 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.004085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.004085Z digest=sha256:ac48aab9aeed656348cfe2b035900b2a1b31dda11a81f05a52f32e1960755c13

Observation ea755578-ed6e-48c1-a8fc-4b270399c075 · outbound

This paper cites Improved baselines with visual instruction tuning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved baselines with visual instruction tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.007607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.007607Z digest=sha256:e0a90b81f1c9e41a79abf0ca7669f037179ec7d78610e97a93a046d4f761462c

Observation c6e4bad6-63ba-4c71-bc04-1cd724b58701 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.011264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.011264Z digest=sha256:7e3bf7386462dac31dfb6f33726187b35d1f3a0560ba70965b718b7490aebcb4

Observation 85856486-bfae-44a9-a157-c0638ef44a41 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.014866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.014866Z digest=sha256:ea5ed5a5aeae9e3a14727c3c6b5defdf7dc463b64aeb116ca18c42086b42f9e4

Observation 866e57e5-febe-4c62-a48f-569875f5fdc9 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.018179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.018179Z digest=sha256:5abd41fe69da0ba1aca83bcf6c21dbf92dd3f152e6eaaf07038c5b3610e8b29c

Observation 4d41cacd-b65d-4fad-ad0e-98d679aad327 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.021821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.021821Z digest=sha256:059f0f227539dd36b0e9cc7ad3636c692df1334a6ed12f71c353ed2444d82a97

Observation 743ef82d-e42c-42e6-9bdd-0f9672ef23e6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.025413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.025413Z digest=sha256:7a567ea59b7b6829af6184c59568fced9ac49687d4c0e403c621fe84a52d0a61

Observation c9f02f92-f888-404a-bb3a-02e0f1357460 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.028712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.028712Z digest=sha256:080ac0764be424d9cb88c75402da09bfeb9635fe2bb27b9f3b6ff80e7fc6220d

Observation a55093e7-af8c-4648-80c8-294a51b3d510 · outbound

This paper cites GPT-4o System Card.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning GPT-4o System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.032184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.032184Z digest=sha256:4a1681b0b784587230db6b58e90b57814c1a6f2385fab2afd12ce7b511ede895

Observation c1f62a4c-2475-4afe-8820-7f6546cb2512 · outbound

This paper cites Claude.https://www.anthropic.com/.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Claude.https://www.anthropic.com/

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.035668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.035668Z digest=sha256:869ed0c04734b71b7439ce409fb40f3f64ce4d13f4d6f955aae0b00ab503a94d

Observation a1f81c7e-41dc-4c9f-a147-ca620750ba25 · outbound

This paper cites Grok.https://x.ai/.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Grok.https://x.ai/

Reference 36

Resolution
parse uncertain
no resolver link, observed 2026-08-07T05:27:05.039295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.039295Z digest=sha256:53a2f8596ddbca112986d94dc77ac0afb83469eee97919e8f265b6370fd24d87

Observation 580d4b2b-7a81-4cf8-a2d8-cf1b67f926ee · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.042756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.042756Z digest=sha256:5479557b4fb65dfd778cac094e352110bc3f1d8d5e908206684af746683d5439

Observation 217880b1-42c6-488a-b0f8-96e668dc1cee · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.046056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.046056Z digest=sha256:16517cade85bd2ca9c7a679f53fbb5ba47760bbb58690623b4b3e6d9dbaa7f87

Observation 373b2e0a-d748-48a8-967e-69b399de2ec1 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.050109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.050109Z digest=sha256:a3a160030df11b359c36354f868ff2c9ede7505bf7dd77259dddde66aec4d25a

Observation f00f1f6c-04fc-4e3f-8008-a8d0088b3fbf · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.053861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.053861Z digest=sha256:ba03804b70a3544c8980d9bac4af326655e90bde5f23cfea5a5a1c2ec3b1eeef

Observation eed55363-e79c-4b63-834e-d9e1eb30dab3 · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.057374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.057374Z digest=sha256:61d77d093f0671b2f90b25a3780dad5040dc366bfa41ad8a230e4ceb988c4010

Observation 2b587f56-f0a4-4f1f-a570-c273cdc5878a · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.060976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.060976Z digest=sha256:96fcc959b757a493c043e54f5fc67ee4eec32210feaaea4feebe98ed854955b8

Observation 5793574b-1f18-435d-8eb9-a6f95939c91f · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.064735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.064735Z digest=sha256:a26889b0cceda5ed3c1d37e864f6fc899b41585a2b03df25d96fe8ee32d3d2cc

Observation 25e96661-8911-472f-b823-13befa0f4b47 · outbound

This paper cites Compositional chain-of- thought prompting for large multimodal models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Compositional chain-of- thought prompting for large multimodal models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.068394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.068394Z digest=sha256:8b96cd8a74f395ade4e7d2befaa4a1cbbb1f6e5bbcf50279b2d7ef75b4db1408

Observation 4093cbf7-d757-4e22-86d6-75ffac09711c · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.072062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.072062Z digest=sha256:8a9bc0ddd03632afe0d94561ace8955f948e46f75d3b2b16f51330f0b1ef0b40

Observation c88c960c-7624-418f-870a-1bde9f3e909d · outbound

This paper cites Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.075376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.075376Z digest=sha256:f795ffae997943c9e5519221ecd770c8fa12e38b36d566faf901e1aeeaf1d963

Observation 722c6d23-e56b-4f1e-8965-a73b0ea02771 · outbound

This paper cites Visual chain-of-thought prompting for knowledge-based visual reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual chain-of-thought prompting for knowledge-based visual reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.078937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.078937Z digest=sha256:9252aab09b38c2cfcf475092a2f2a9398fab6a52e29d55010f8e627b5a65872a

Observation 8b2136dc-12f3-45aa-8bda-ac0ed1d08a1c · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.083376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.083376Z digest=sha256:1bcaa0a756c1b237de98808faf9c559d08840597bd8b0d1a855e6ffad3c26a66

Observation b21cc11a-7cb7-4090-b1d7-b43f9e8a02c6 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.088125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.088125Z digest=sha256:75f66c260dcebf32c40dd78770b1dfa3cb876c627da93fdf82efaaea31805e4a

Observation 38fca845-86fc-4356-abf4-22a29e52629f · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.091700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.091700Z digest=sha256:a75b4b0e6d325abad0e36bcc4661cb9898c43287d08c98a4fa0101be330388ca

Observation c4e8ca4e-e792-4925-8635-57e9b487c9aa · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.095888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.095888Z digest=sha256:dd3fbc528b155cd75608e3f0dc23a6de4844232f746f1bb5d0897b22630040da

Observation d226f46e-6381-44ed-a5dd-84a08c35a2d9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.099935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.099935Z digest=sha256:e54a3adef73380ef6b7b275eca4e9275f599eb4e76b5a493db6591f9da8975ce

Observation f129975a-ae54-4a41-a608-82aca32d2b88 · outbound

This paper cites Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.arXiv preprint arXiv:2504.14642, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.arXiv preprint arXiv:2504.14642, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.103227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.103227Z digest=sha256:377662384bb39524a52e2eaed7cee85259568132b72bfbcdce2b69dd1f0308e4

Observation aa3c6233-1240-410d-b735-770dffa707f5 · outbound

This paper cites Compile Scene Graphs with Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Compile Scene Graphs with Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.106433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.106433Z digest=sha256:bd79f74b7fdb5032b524098d91ebf7b724257881bac6257f8085d914e6532faf

Observation fb0606e7-931c-4f6b-a038-f2783832d086 · outbound

This paper cites Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.110383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.110383Z digest=sha256:4b7d618834aa4578d1085ba78ecff71f00ed66962b7190ec8cc20a26194c92e7

Observation 900592f3-cb75-485e-8f6d-57cc9329113b · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.114508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.114508Z digest=sha256:bd0f5ab48b71a7db9c72dccad20237219e3a418d11a2831a0a7500ddaa279495

Observation d4216eaa-f78b-4267-8cf5-86dd6946c4ec · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.118106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.118106Z digest=sha256:d97d22cfe3bffba75ced92b9342717fd82ab2731dc0f1ee858ca5d8231e7a846

Observation 103c64de-bd83-4be8-9b3f-cd361c7ee95f · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.122122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.122122Z digest=sha256:a6bc8fe54888049153cf28e6e0581ed639c7eaad182c2fa313bf46ee8e37777a

Observation d120e7ae-903b-489a-aa89-14f38d5d8a2f · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.125750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.125750Z digest=sha256:e4b0423a3f62f5e560d9b7fa428f2b23eb5b32950119dbe7a267f1f572bcc3d1

Observation 37fd741a-0b79-46df-9e00-a1c67bf22259 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.129486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.129486Z digest=sha256:307de4938ecceaf9826e7dd06ad8b6e8ca771b9e126aabcc6e2f1115fd140892

Observation 7679a474-d1ab-4409-abd9-9c4aaa7000b6 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.133072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.133072Z digest=sha256:053aa3d1d437fef86cb6b85329669a4075d509946ec4af4b7d4900aae4ff10eb

Observation 6348801c-8c22-4cb9-8943-fd8dbb1ab974 · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.136689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.136689Z digest=sha256:f0f5fdbeb6f3761311755d891c3c273b32ea061f4fae56b844b339a0d3dd1fca

Observation 22bc375c-4e64-45da-a171-f57595902bb2 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.140729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.140729Z digest=sha256:a1f518284b052a3dc0bc98ff452b6937901fd20d0a8ccde2636f2d6be6b4eb28

Observation d157d2d5-3671-4249-8e5b-5e86a6654f91 · outbound

This paper cites Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.arXiv preprint arXiv:2504.03724, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.arXiv preprint arXiv:2504.03724, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.145217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.145217Z digest=sha256:7614073b93f3bc0cf4a91c0b18668427d45ff01029a1e55c88616d8d6339c1d6

Observation 4217f07e-78ef-41aa-84e8-8aa534724142 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.148735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.148735Z digest=sha256:b523f299ee425184264784cec50391c5fed526baba99e02b48198bd06332fe63

Observation 5fc15cf8-7aa9-4a9e-8c30-3b2432984ada · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.152302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.152302Z digest=sha256:bc7dbd15377baf40a14dcff8600a7f7cd66bb1a72de700f2a6b152d41de93a76

Observation c581d90e-8ee5-493f-9b53-e0e95a7b9f11 · outbound

This paper cites Microsoft coco: Common objects in context.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Microsoft coco: Common objects in context

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.156233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.156233Z digest=sha256:f51cc634e39771868e82353a1fcba85956ac47d5dfc7a869607e485d9263f50b

Observation 2d04fc18-ed40-4c4a-a6e7-95235979c46f · outbound

This paper cites Segment anything.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Segment anything

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.160052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.160052Z digest=sha256:503523802037a4c85dae071669f1de61d89a9c0bcaa40a1065fb1e08f824107a

Observation 89861dc0-b9bf-4763-a63b-48c10926ab6c · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.164402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.164402Z digest=sha256:bf2c8a9510b8d5f887f544ef70799cfe9672bd1f89e325df5748d5a943d2e2a0

Observation e381a6a3-9853-40d7-8bf7-9d101078f2ce · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.168039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.168039Z digest=sha256:8684c67cd33754862430d4359dafe15dfad04b347f1326bf0cde7ed73ce13cdb

Observation 8f3867a4-b03c-4462-b451-93b616482945 · outbound

This paper cites Dual-glance model for deciphering social relationships.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Dual-glance model for deciphering social relationships

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.171480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.171480Z digest=sha256:913f008692c21348dd4ed0724581c07268e75c87d5bc164cc11c4a8f3304e24e

Observation ca7b2491-d32b-4813-98bb-6b7c5dc20abb · outbound

This paper cites Towards vqa models that can read.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Towards vqa models that can read

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.175095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.175095Z digest=sha256:46d8808db5d3f61f11c6c60a233eef99929ee4ede89c630026004a110625a540

Observation 291d4e8a-b359-415c-b459-b6c1c0a85b02 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Docvqa: A dataset for vqa on document images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.178737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.178737Z digest=sha256:2146316c2c4f4be44ddc3897367c5f5bd685290ded29058bdfd5bc876d8d8515

Observation a10dcfd0-96a0-4c4e-8c9c-f3242c3acade · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ocr-vqa: Visual question answering by reading text in images

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.181987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.181987Z digest=sha256:33871fbbf187f75184d05fe6b6a0c1444c110a6420b3fc0de7b8cee58cb5d27a

Observation 068b58b0-13bb-4fbd-9d9e-84d6034677ae · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.185608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.185608Z digest=sha256:1239bb859486f435739a1c7e6f972a726826dff9185c09c4d4acd0cd329ade02

Observation b9aa5595-ee68-44df-bb08-01cfb3932d8d · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.189776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.189776Z digest=sha256:f0bb4f21fcec8a39bdb0042e6efc40d22bfd8f2c633cf852a793957ff8f07e42

Observation 7d62ce74-5137-4911-8044-048da7265ebd · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.193367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.193367Z digest=sha256:f7e4ba003f67ed609d0dc450a035f728a184262c7fe034293ad05a1427dec207

Observation b06efcd8-19f4-4144-843f-9821c78a36cc · outbound

This paper cites A diagram is worth a dozen images.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning A diagram is worth a dozen images

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.196950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.196950Z digest=sha256:2191207edbc18d780df0ba1ca1dbd5b114a9d74a9a0b7181bd9e4d8db8ba6082

Observation eabeb259-7914-4223-8ee5-9a69dae88dbc · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.200325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.200325Z digest=sha256:e35036525f996679b02bedcca9dc25d354d071f6ef05dea55c7b7007bb79307d

Observation 6cca9462-d10f-4096-afbf-ed6e80f2ff0e · outbound

This paper cites Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.204087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.204087Z digest=sha256:53e78380f3cd4a7ef16693d4fbddc49c713ec68cb27911113d417533434aca14

Observation de39bb08-53f5-4639-9513-e13008944536 · outbound

This paper cites Quality at a glance: An audit of web-crawled multilingual datasets.Transactions of the Association for Computational Linguistics, 10:50–72, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Quality at a glance: An audit of web-crawled multilingual datasets.Transactions of the Association for Computational Linguistics, 10:50–72, 2022

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.207819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.207819Z digest=sha256:598a4dd09adea2a5aebd0ec4a99d13391a07b113d6ffe528f9edc88a84ef4a1c

Observation 453ec9e4-fd2a-452c-b484-b24161f7711f · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.211152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.211152Z digest=sha256:379e7f59d8627839e10e7511319720148d059bb66848d0d5a4a4bb91878e6f65

Observation 2217614f-6fb3-46a2-bb53-017653874857 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.214816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.214816Z digest=sha256:e03e0e95252bb1f9e9db864e6d165a786e408f6934b7687fedecb9e38e2814b0

Observation 31eea80b-7b86-4b09-a24c-f21b2b5d36fa · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Measuring multimodal mathematical reasoning with math-vision dataset

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.218695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.218695Z digest=sha256:48f88667b2b15b8d930404f4a7902e28955a79a53e205706deccfc1a8f64913a

Observation dc998b26-193b-44a7-8a6f-d61669a8b7a9 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.221981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.221981Z digest=sha256:ca936167601a045c12eb7d5882bcbfec691c857ab82166f0d6ccc74f6bd0583e

Observation a5bc9403-a0d6-4efd-b92f-bc0e7a252ec1 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.225253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.225253Z digest=sha256:d83621e992b8f1d598642ec29d51a7f42e2209365dccb9920094c803a71fdf8e

Observation d11d3584-06d5-4374-867b-ff60a6a0b96f · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.229610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.229610Z digest=sha256:89c28cf52dc13b999b31be58939bd34c085d89ac83a1a37f63f921b1c7d9e7d1

Observation 2583034b-a82f-497d-8e02-8cfb62091901 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.233360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.233360Z digest=sha256:8225b14432a86ddf1b1f250cef1893c5d0bf729cd47471f3f1d25781dfd833dd

Observation 7e6b901e-2631-4aa3-8984-2716b6426aa0 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.237152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.237152Z digest=sha256:d63607d86a8212be51463f7775cab818f1673d0b604a543f4b444d35dd067c7c

Observation 6017241e-0ff7-461c-8c6f-b5ba0517dc04 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.240629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.240629Z digest=sha256:8d633ef1e3469774dfcd9abb3a61610de3e74112d6127df080babfb5591c3444

Observation c3dbcfa2-99e2-44f3-b95c-5f8a0ae48840 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.244024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.244024Z digest=sha256:154be9d2a574ebd09a1f2ff97ee57bcca618b156416b462f27479994e19469b4

Observation 98464604-acdf-499c-acf3-4e38f8e7e286 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.248472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.248472Z digest=sha256:1677357946e2d5d4e7e9c3c07bd4a6e8e685c50341939ac8cb5770ebb4cdafd3

Observation f24ffa82-7d1f-4a5a-818f-553dd3864afa · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.253190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.253190Z digest=sha256:552bc35cb7f9a1f7815f49f5f750b40cd850bae0e4d77eaefb2f2e3baf8c33d3

Observation 00d2368b-9ee9-4596-bafa-2c10545d8f6f · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.256916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.256916Z digest=sha256:39adf38db30c6ceef7e0a7eaf6996bf4fe081f5140749776699d65b15da8cedd

Observation fed61cc2-61ea-49e4-bd12-eddbad3687f9 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.261384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.261384Z digest=sha256:8f5ad6b65638002e44d38e99a5da596b4e289f3f6478911646a20fe9d73b5389

Observation 2c09f3aa-a7bc-4afb-93a0-874cd4d96168 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.265419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.265419Z digest=sha256:0a4bd9216d240375f2e225e27e738350c08d3af924d558f2a245fc57e7bf5ff2

Observation 0738cffd-d5a9-412b-a106-c939684827a1 · outbound

This paper cites DeepSeek-V3 Technical Report.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeek-V3 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.269421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.269421Z digest=sha256:b2848ab8b4ee7b59d5511eef6e9f79d62b765d513be950b49d2bf140549ea4b3

Observation dcc8fc6f-9d29-414e-90d5-719c3458ef99 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.273613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.273613Z digest=sha256:40a8db095e259e8b117f44febb682b87cb48c7908e34566a69844648685c6198

Observation 40909008-8a02-4e26-98bd-9a1ed8bbe2eb · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.277827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.277827Z digest=sha256:4ab2927c1e1270cf07b9be856dd00bebca83ddfbe9817334bbec3382547152cc

Observation 522bca22-8ced-4f8f-b805-8fd38471bdb4 · outbound

This paper cites an unresolved cited work.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.282142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.282142Z digest=sha256:d06968253d43f0d5a8c98dfc3d30f9571eea24882cbf3638ae3a9b55200704aa

Pith citing papers

Observation db245629-0bcf-49c4-9b7c-5aa648952409 · inbound

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention cites this paper.

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:07.475851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:07.475851Z digest=sha256:4ccc3b99e28505e28e08974e394b30255be3f446872341183830783462d720a6

Observation 81f9e82d-4567-4976-9de0-b0ba9a4a3bb7 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.158919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.158919Z digest=sha256:62d8611d26caadfbee71386e76a363b2ac90a6237f1aacf184d07ed5ffb3ab34

Observation 48e25eed-7b27-47cf-834e-9b44876c6c93 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.853224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:bcae7b44d9b306249326218dcf8effeb955a129eb22b4fa5edddef98cd2804a8

Observation 0664b863-da57-4521-b0a4-9e6706ad160d · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.176406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.176406Z digest=sha256:46311352a5da8e9d276b1636764555509f6c2d9e3be3ef135913ab2d4f8013f7

Observation c6b81be2-5b5c-4e7c-a979-e9098ce6c857 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:07.983014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:b1c2144099bb82e3fec96f785a2f6b7287b146bb69ddcad3c32515bc1bf75a58

Observation 6983eb48-fcc7-4ed4-9e77-1de6f3922bfd · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.563530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:8c0ee3a7eea01e5efe825d11095b0de7ee2bbb59a4f1fc1326f36eecb30304a5

Observation 09b8ffd8-cfb9-4af7-b9d0-3ed26b71c759 · inbound

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct cites this paper.

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:46.106878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:34:27.719022Z digest=sha256:670035f24dec750cfe58b1b85b77e11132ec0eba67cb0f850177a7dc0c997baa

Observation a38f450c-7479-419e-a93a-2d14d92ed709 · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.164827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.164827Z digest=sha256:d96bf9a036c5b55654bdd691ee15587a412ab91b00dcc5738d236c6b7c9f2f62