Pith. sign in

Paper Citation Record · LEDGER

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

As of 9 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 6 inbound Pith citation observations for arXiv:2506.01078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01078 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.498636Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:37:19.575834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.331710Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa0bce29-ef9c-4bd5-901f-3e684ab67a49 · outbound

This paper cites GPT-4 Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.774631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.774631Z digest=sha256:638e0ef9ab01a6d02ab7ca2c89a34938d127887040024117c06164950172c18a

Observation 12445f70-6053-460c-9e82-f2362571e48e · outbound

This paper cites Qwen2.5-VL Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.821566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.821566Z digest=sha256:8a83b7c277d8f5a2a624e864a74089be34dd68c8fcc55847a1e159dea4b287c2

Observation 38f64302-aa0c-470f-99a1-8791e16daa41 · outbound

This paper cites Perception Tokens Enhance Visual Reasoning in Multimodal Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.856611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.856611Z digest=sha256:dd80ca924794bea3e79a15b13cf4a88be4f94be4b51343b9d370b825f42cbe86

Observation 9ecec8f3-af2c-44e1-964e-f995accd4dbc · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.906301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.906301Z digest=sha256:b66a9ba852403cb5234474d99dea823a778cbebe62e69f6b81bb49f30ba9cea7

Observation 023d15c5-a23f-41eb-9564-d013bc58a441 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.930951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.930951Z digest=sha256:75202a8e03ff42d0a128cb27bc681eb4f17a53a32eafa7e152614bc08c7ef212

Observation ebac17ce-1efc-4864-8848-c04bf13502bd · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.969152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.969152Z digest=sha256:66c227692189d6e5f12c0c7fb92c06374eb580dbe4fb9edec859b2730e1acf62

Observation c9ff8af9-fef5-462a-8545-4a856cbd90e8 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.014182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.014182Z digest=sha256:d78b074d26954c069cfff9c612cc96675b4966650de5f9e78e820a0d835b8173

Observation b0a05cc2-2c03-4da4-bf31-5f88bf5219d4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.061180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.061180Z digest=sha256:cf6d26d3ca973f5520169b20b26e2e937f20d63ddcd25b2487f1122178dc191c

Observation 4d17a14d-d0f8-4327-9080-26d7492ad706 · outbound

This paper cites Gemini 2.5 pro preview model card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Gemini 2.5 pro preview model card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.108675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.108675Z digest=sha256:c7e11a700f4b4cc2eabcdded8ddae124d2bfd1c97f48028f47f39716528139e2

Observation f1e0314c-df3b-4a52-8752-6ab3de98c9fb · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.149124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.149124Z digest=sha256:02d96a64e4cc3514af1e4a0263b9295ddcb554586e2472116df50f4597d58175

Observation d879ea55-65b0-4411-9b04-86a11e4af277 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.213124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.213124Z digest=sha256:25a9aefdab19ac125c349530d48944f50b7fcdeea6424cf93fac6c9575d2822f

Observation c25cc8e7-7529-47dc-9568-0ef4fd50a058 · outbound

This paper cites Cantor: Inspiring multimodal chain-of-thought of mllm.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Cantor: Inspiring multimodal chain-of-thought of mllm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.465045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:12.262102Z digest=sha256:fca33222865f569018027f0fac80cb71f5a2e99363975f74e35660d9632ac816

Observation f24ee158-fd32-435c-8c58-fa50cdb828ab · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.301566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.301566Z digest=sha256:0a780fa5b3a51a97dcceb19220281ccfc06a49bad1f88fe4e7d14fe0128c0793

Observation 5021add6-3abe-4e19-9a91-c1f21a66fc32 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Measuring mathematical problem solving with the math dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.290620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:12.345211Z digest=sha256:c553e8899d0f22846c7fd5eae50e7aaa807d16c701e77c46de40a17d27c52bd5

Observation 7efaa70e-e358-41cc-8c91-60c873221cd6 · outbound

This paper cites The abduction of sherlock holmes: A dataset for visual abductive reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking The abduction of sherlock holmes: A dataset for visual abductive reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.168420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:12.398632Z digest=sha256:691a1dc0a606b860864fa5679404b5fd058f12765f96715f337a31472e811477

Observation f35b0048-75c2-496e-b003-074ef2392571 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.450962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.450962Z digest=sha256:ee42266687eb22cb17a8ee80e5a22ffa4757dda7affbeb0adaadf5f200a6e815

Observation e41420bb-eee7-4219-a427-1d2970955726 · outbound

This paper cites GPT-4o System Card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.506068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.506068Z digest=sha256:c8a37333d9a245cc86fb80be6e238a38c166c6d00a9d7cd7a91f0bbd66528be1

Observation 11301c7e-0e1d-476e-9355-b371fd77dd1b · outbound

This paper cites Math-Verify: Math Verification Library, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Math-Verify: Math Verification Library, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.006825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:12.549003Z digest=sha256:a0f95c57910b7aded43bc4b7379b2c285a4a6cf73ecef77727499fa462d3093c

Observation 7e1d0daf-db50-469e-be9e-1c0f31cb1850 · outbound

This paper cites OpenAI o1 System Card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.601005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.601005Z digest=sha256:6c606451fedd1d19dfcefae01500dec6759b19f19500403abbb53c5fcecbd863

Observation 46c18466-ed5b-4404-b7fc-2097d1c65101 · outbound

This paper cites Abstract visual reasoning with tangram shapes.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Abstract visual reasoning with tangram shapes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.814552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:12.636382Z digest=sha256:9a37cec3d7a39e056e4cbd00a153455b39e640fbf3de594f2d98273732f27701

Observation e0ac755b-9f00-462d-ab09-3b38b42ce4ec · outbound

This paper cites Dcot: Dual chain-of-thought prompting for large multimodal models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Dcot: Dual chain-of-thought prompting for large multimodal models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.642575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:12.689346Z digest=sha256:2940f28ac6826a1041b42a83d95136193dcdaaf30ce4f6c7e8d4104763403092

Observation 3b2c1749-b529-46e6-b946-27af2bc38d7e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.734446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.734446Z digest=sha256:9407bd9071f92e9d0839a6ba47042f4d6862e8f7a75b223ea1a34752967a3d96

Observation 0eb33f72-d27a-4a76-930a-3606de24e0fc · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Silkie: Preference Distillation for Large Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.792022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.792022Z digest=sha256:ca7ea72aae27646a695ca8d292a388532940e626668526985dc1c5c6af04170a

Observation 62f27d10-92d2-4000-9c11-92909d519264 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.887272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.887272Z digest=sha256:98a92c860caab83bc6b087dea567cbd60e76fae9d58763abef9bf0db810cda28

Observation 1208e7d4-dfad-4aa9-ba0d-592b73ae6df9 · outbound

This paper cites Diving into Self-Evolving Training for Multimodal Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Diving into Self-Evolving Training for Multimodal Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.938364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.938364Z digest=sha256:8619c5795a33031758592b8032db112b509828cbec674218f0f1ba3c0150d77d

Observation 3194040f-a74e-4226-818d-3234125cc836 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.978031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.978031Z digest=sha256:93f9a91d512eb4665aa811d18689919365e6d70c8f97f7a87e5e4037c953d726

Observation d0bb0f6b-b5c6-47ba-b795-aefcf70039dd · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.510804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:13.031238Z digest=sha256:65a2488622821a4cac2e1d3221b4b368320c51bf808093c5957940e2f10084a1

Observation e43b2e78-e870-4fba-83c5-a9fc52b45e50 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.090908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.090908Z digest=sha256:e9fbfcfdae779e588e12e837d88128119596e789c158b502c9da225004156109

Observation baf9e354-97f7-452c-aeb2-9d34a4aef0c0 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.438773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:13.159210Z digest=sha256:d1df47d52229262eea8c30e2ac3c7d0d98e509c63f01f9653e3bb6c81fcc0fd5

Observation f3ea50c4-9b66-4532-954b-befd96fe0118 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.202412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.202412Z digest=sha256:efe8dccf136d82b57d55bbb1d52aa199ca86189af4bb5251e572e0c0b143cb27

Observation 768f4ee9-819d-409a-82a3-6511bcca86bf · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.230798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.230798Z digest=sha256:ed695646e72662037860cc5f3f9dd14e07bff0560c744ee5c0e0bc77e5922a8c

Observation 85d412e4-cd40-4c36-8b56-4669993426c0 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.261487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.261487Z digest=sha256:6d95495dd68c702272b7d77e1668116e197bea85a592279f06622e83ed3f9850

Observation 86254fe8-b730-415d-851c-388c2eae1875 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Compositional chain-of-thought prompting for large multimodal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.282182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:13.311730Z digest=sha256:a4507586d289910435488e40e6a2756d68df6121008ac1b726d6d5fc8aa94076

Observation d1e06d9c-a647-4703-aa0f-1712997adb57 · outbound

This paper cites O3 and o4-mini system card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking O3 and o4-mini system card

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.169036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:13.342588Z digest=sha256:1145f9112c27bf48998fac4ecc7aeef16978c74a0ec9f50b6f5ebaebd447d26d

Observation d2cd426f-ae60-4c5b-a659-a910706b907c · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.380294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.380294Z digest=sha256:d2f61e2cdd41020b890326dbcc38b6a898c70d484574f33e6848e7170dbf5c25

Observation 7f4e8530-e3d7-4994-a8a2-82b77e6730e9 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.434001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.434001Z digest=sha256:7cd647c0c38ea717c2bd47c47b4a8612316d963ceea5cdffb3132f6b4bfa14a0

Observation 7b8c2e92-d821-40c1-b36f-71bedf13125f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.463903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.463903Z digest=sha256:0e01839671cd2268dca0de24c63513370707d8d9a75a2402261ba8cb24c99787

Observation c6d3d20f-5991-4e78-8cee-5c35904d0653 · outbound

This paper cites Rethinking Reflection in Pre-Training.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Rethinking Reflection in Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.515733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.515733Z digest=sha256:f94bc71475b2e4533ec8a3fe223be12caae06dab6fb4d136e12f91e0c630c0a9

Observation 9ab4e716-36fc-488d-9445-0b2094067eea · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.537018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.537018Z digest=sha256:28465464acc356ba4e18b894b10cbdff63453a402805c9af7b284783e1ec7322

Observation c9f8922b-8b8f-40fb-ad7f-a5e4a598770b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.586726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.586726Z digest=sha256:937084a64c95087590497530c554bdc6dd057a33aa6a875b1c6c7491a04acd00

Observation 983461eb-e95b-4f81-bc9b-e7161ec4e2e9 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.667498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.667498Z digest=sha256:12078c0774511acbcbaff6bac2e53c5355c124e36bee335a93d99e273568e027

Observation 9623ce81-8064-42c0-b353-2e803c167ba6 · outbound

This paper cites Kimi-VL Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Kimi-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.706010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.706010Z digest=sha256:71133d2b09ffc9aa9ec411f8f63f4161a2bf417437bf35694c26ccf210fac80f

Observation 443fff95-c174-4f1b-b701-1aa212e2e4d1 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2025.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Qwq-32b: Embracing the power of reinforcement learning, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.751900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.751900Z digest=sha256:108b06d0eac0be216b89cb305a34acb72329fa5a1cebe8a4f3d85d7beac86eb9

Observation cf62bded-4187-46b7-9d7e-cff9c6edfacc · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.783825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.783825Z digest=sha256:9fd96059f116736894da9a4d7e25be419941d38e54f5f10d064875f9d8cfa39e

Observation 9a19dd4a-05ab-423d-8799-5b2d39bfd03e · outbound

This paper cites Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.815176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.815176Z digest=sha256:32e846f444b2867899ef2e96713eef46276f4b94267d0ae26a33d0a29e8b99b3

Observation 7e2bbec1-3765-486d-be83-792e41dc3a9c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.847422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.847422Z digest=sha256:9c5088c0692bdfee74b89d534c6b3fc3fabb1aa1cd3fd06b1b9d5492787fb2e8

Observation db4734aa-bd8f-48fb-8b6d-ad871887ec71 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.922969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.922969Z digest=sha256:5c5fe11bfc2198ca3c2cde4dc28142eeb433eed0739c27d3c3817a8bd1a883b9

Observation a17f981c-2931-4cd4-a188-00780aeb2972 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.976791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.976791Z digest=sha256:eac2bdc97a3d8826ab5c382f07a8f5b79644922abb8588a360de317fec69a32a

Observation af52cd2a-3648-4a19-8272-91c349c85350 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.016315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.016315Z digest=sha256:4f8db035335ee89b56b307c18ef12b576707f83122aa5c3ea40207ef5e7038e7

Observation be5d73bb-39cd-4d6c-9ab7-4903ac947a3b · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking V?: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.047663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.047663Z digest=sha256:3ff68618af6a5887fe7b3caed39a61b93bb2f9a419a1b7cbc013c91ccab5a47c

Observation 6f645a90-1413-4158-bc97-bd2620bea1e5 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.082910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.082910Z digest=sha256:d775a56aba5c0923a7df8e332324b17422cff09ca4be72b0a961771c28ea740c

Observation f7ffcbb5-699a-45e6-a7b4-69e78ef5f9e6 · outbound

This paper cites Valley2: Exploring Multimodal Models with Scalable Vision-Language Design.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.123082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.123082Z digest=sha256:7acbf5e187d4aae493ea03fedcbab34486ec668c028ef12f7a9f52af6fd1ca1c

Observation 325687c6-cd93-4575-9e89-36946bb2e03e · outbound

This paper cites Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.044283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:14.154557Z digest=sha256:3f99587acdb04a04eda59343fbe805f6a24a088688a7f5f7febc3e9b3e6a4217

Observation 8216b34c-7b24-4c7d-b699-180713ba760f · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.196592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.196592Z digest=sha256:d28eeec6e282c75f40a056952719cd400435c7fc4eac6519afd12f23e3ea3456

Observation ef1a0837-1386-407b-829c-63f946f193cb · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.243341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.243341Z digest=sha256:f821f6c2b624fe93e5487fe7dd5ffc55ab56133e5367dcf050e97a6e0d00d85e

Observation 126c7a05-0e26-4e47-af71-db622ff377b7 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.280040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.280040Z digest=sha256:7f140c60ddebfcd3f839c0be5d20503ffa606c6611cf18fff1200d5d8864c634

Observation 5cb9a2fa-18d6-447e-a62d-cf036c3f1d21 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.320981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.320981Z digest=sha256:a7bae6824f69b31d3ca4a226ac667556d0694402f92bd95c7365cc26c1c7fa3c

Observation f2a6f081-263a-452c-8fcf-e02f9268efb1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.363044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.363044Z digest=sha256:191e66a1a3316b21df3afe2def0b1c480d53ec5329bfec1022085eafa583210d

Observation be287d4e-fbca-4dcb-bccf-494b8b5b684d · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.405560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.405560Z digest=sha256:a0e6a0b9089909455a52038c89e6f0129c9f754902a120fca4fee9bf52c84530

Observation eba62b12-8943-4eb0-90cc-1e83da13778c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.467663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.467663Z digest=sha256:df27725df8bfce9c99a9395a6aa4bc543217acc696d76764834f9a7b60ad87ce

Observation 5fc1b220-d50f-4c6e-8177-b57b88ed8c5e · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Griffon: Spelling out all object locations at any granularity with large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.944518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:14.511694Z digest=sha256:9b5c7177c037cabb1fdf4d425d84a2df00a0fd601a3c5608e3e3d5e42747f5e3

Observation 2d6a7a05-c0da-436e-bf2d-44ed9fd03936 · outbound

This paper cites Ferret-v2: An improved baseline for referring and grounding with large language models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ferret-v2: An improved baseline for referring and grounding with large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.766547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:14.559871Z digest=sha256:0f2743f8327ec75292d32da3bfad14af3d0366f1e54a1c15917b54ef88aa96b6

Observation 9d312232-0a92-41c3-a7d8-a8b6e7113e02 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Improve Vision Language Model Chain-of-thought Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.605390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.605390Z digest=sha256:e0878aa8605a4f56d4bdcfd78b20b933e53a789cf238d1637d9bf79c6091dcc8

Observation 4a5719e8-17ef-4d2f-994f-0806687c0102 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.660602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.660602Z digest=sha256:c45e27e77579488e345df0c3985042484e224299154e25289d46a0a48486cbff

Observation af67a684-595c-4119-b025-89b1851ed955 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.702270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.702270Z digest=sha256:ce1bf94d5e0ff1676847d2e8411ba60f960f4f8f410b3cd2ebc89feeb36ba383

Observation 5e94eb29-f3b2-481f-9344-cd4227812f63 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Multimodal Chain-of-Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.744601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.744601Z digest=sha256:5783ee79078485c8716d16fe3a14a7201524900d67e4d30a598c1180428419a9

Observation b61ee54b-df04-4705-99f2-a932e7f313c4 · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.808518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.808518Z digest=sha256:cdee80bda69a32c4cda6ba611ac01e47068b56b53ad4c1a5b5384ddf36f1fd01

Observation ab1a015a-c659-4ae2-979b-d101bc05bf37 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.847482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.847482Z digest=sha256:a9915b47440fe3036918e59b93cb472ab388c9e432dfb9a9e44fd8009cf58083

Observation 154118fb-2756-4d03-a647-3194e61ddeed · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.641879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:14.880283Z digest=sha256:cd887cef06ea9ee01e67fc8ce1ac8b77980c3e70f83cfb6372c1d6b30bd7073d

Observation 49b668c6-1839-40b1-b341-9a14a4da1826 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.545042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:14.915753Z digest=sha256:d062cbb6ff2d3c07dbc71401ce2ccad3e62946bb8a66ec1916cf1d12efbbcc3f

Observation 8b139858-c0d0-47db-92a4-2475a2512a97 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.460148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:14.970907Z digest=sha256:3ab568274f7ec63b82f74b5697ad1f9c5aada873aaef02f3a409a63e61c618b7

Observation 3bb65ca0-0f86-4c04-9581-9da1db0c5bde · outbound

This paper cites model’s chain-of-thought (CoT).

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking model’s chain-of-thought (CoT)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.375472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.015345Z digest=sha256:023e814bd99605f3560656a00ef76b751a9f3a6b9a9c01ba2487e0c12221baf5

Observation 0337b9b7-4972-4951-80b1-eb404c4cef5c · outbound

This paper cites - Then, wrap the model’s entire thought process in <think></think>.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - Then, wrap the model’s entire thought process in <think></think>

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.288363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.075607Z digest=sha256:9c758b510d591da8c69f9b7425713915acafe2f61e4e655f34b897c3d9ec9d54

Observation 03b5c2e2-ab2a-49f7-85a9-c3c61f023780 · outbound

This paper cites <vcues_1>, <vcues_2>,.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking <vcues_1>, <vcues_2>,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.211611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.125236Z digest=sha256:e055a97634fd0f65825d667e0e3c0f25ef5c5e6211f1c5903969dd87ec1c140c

Observation 92a74542-590d-4e50-a054-f3ad6876cce2 · outbound

This paper cites All the data to be processed now concern reasoning errors based on visual cues rather than errors in visual cue perception.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking All the data to be processed now concern reasoning errors based on visual cues rather than errors in visual cue perception

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.139294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.156197Z digest=sha256:a6f227f2e44f3d895bf722cc498b68b5d4f6fd0919b6771ad17fb25b258f1679

Observation b2715fd1-294c-463c-987c-a6d13ec84414 · outbound

This paper cites based on the rationale.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking based on the rationale

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.067581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.182656Z digest=sha256:ee78b3f796be9005cabd59e22886afbe82b6c4d33e0fa830ea21e792c362a8be

Observation 9abc4366-1090-4d11-b7c8-1a1192610bac · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.981376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.223161Z digest=sha256:b1a761adc5def2838b0069b97392a719e06326e276ecaaf139b8edf4e49a59eb

Observation b3501fb9-f946-451a-ae49-8aa9c067f709 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.906529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.258999Z digest=sha256:1d8d1edff80777ebb1ce87d02a250416f48f124d398aca4b3f1e2003495aa245

Observation 0ec0a9da-6c2b-42c2-82e7-bd83727448e0 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.830330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.298530Z digest=sha256:31b3cbfede2ee3cf9e67f5663c4eba1896abe8d5abe6a91f2ea77a8b09dda26f

Observation 7efba94b-a213-4ed4-a25a-0e79e112b29f · outbound

This paper cites Let's verify each visual cue and its reasoning before finalizing the answer.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Let's verify each visual cue and its reasoning before finalizing the answer

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.740232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.335118Z digest=sha256:99bf9355e4ac38dc71961a32654a4aea831d7cc983c341790e1ee1451cde4845

Observation 37267687-4afd-40a6-98bb-e09353490526 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.651087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.400221Z digest=sha256:987d3220482f0d96b9923d6b6fc87d08fc045beca11384ba5030fc92ae11e54f

Observation 8c3afcc1-d706-4bba-8d3f-c0c4259cfd39 · outbound

This paper cites - The angle 78° is an interior angle of the triangle, and angle 1 is 42°.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - The angle 78° is an interior angle of the triangle, and angle 1 is 42°

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.564150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.434526Z digest=sha256:3a4a70fe13d8dc876c26db46a59c90134fc7bdb16c6ff282f811b92362d3dcec

Observation 1993d855-5ca4-4b5a-a7dc-9b2a6750ef4a · outbound

This paper cites However, upon reevaluating the problem, it appears there might be a misunderstanding in the interpretation of the angles.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking However, upon reevaluating the problem, it appears there might be a misunderstanding in the interpretation of the angles

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.470842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.463161Z digest=sha256:c064e5b4d4149c27481b78a93d6a4e8189e91916fdc0e75a0cc957b17d6fee8b

Observation 27496727-abc3-4c17-824f-ace683444b01 · outbound

This paper cites - <vcues_2>Angle 1 is 42°</vcues_2>.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - <vcues_2>Angle 1 is 42°</vcues_2>

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.279875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:55:15.498636Z digest=sha256:56b07e09004b784712d5bf192913064d2e6f15da91951c968c4671b8c6c97c12

Pith citing papers

Observation 9bf62d37-bacf-46c9-9528-1e292c4f2fe2 · inbound

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law cites this paper.

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.575834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.575834Z digest=sha256:d166a4632ab51818a754bd1c70898d211856ab0c103f6923faec125dcadb1c9c

Observation 51434a38-27d4-4cdd-9a2e-fd34b8e95e25 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.600332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:2164ecb1685619385923de80fa619844a4d91d0283f497ef25cf4cd614ad3836

Observation 74e75a2a-7660-4070-9f2a-35d038cf8ac1 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:46.730039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:46.730039Z digest=sha256:9bf65a50de7b79c8ffa498c0b0d9374e43d7afb9da66318a12d2610e92cc940a

Observation 3b5c11b0-909c-4144-9dbe-7e7e7e6e00a9 · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.576402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:1b8da783e9877e36a8280b975c420416f14480ab3fd114583db7df298f7c4590

Observation 8160789e-f389-48e2-ab36-3b6de0efeffc · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.576609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:d9c208cd0ffe6e1cc2f9b7bfd88bb54ffd9e2adc1512b4dbc3eee83d65ec84bf

Observation 4ea4b481-5c01-42f1-8ad5-12453ff09425 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.333992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:8a9fb44f0f69394009d2c5077783bceb11eb9fd485f6a96c9280de59590c45b2