Pith. sign in

Paper Citation Record · LEDGER

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

As of 14 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 6 inbound Pith citation observations for arXiv:2506.01078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01078 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.498636Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:37:19.575834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.331710Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa0bce29-ef9c-4bd5-901f-3e684ab67a49 · outbound

This paper cites GPT-4 Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.774631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.774631Z digest=sha256:b12504638a4db8655f5e21953232d0945b6a273c2039352d91ffa25035eba5e4

Observation 12445f70-6053-460c-9e82-f2362571e48e · outbound

This paper cites Qwen2.5-VL Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.821566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.821566Z digest=sha256:978962be764402c49aecaf4e4e720f57567837ed5dad01a792e96edad1a16690

Observation 38f64302-aa0c-470f-99a1-8791e16daa41 · outbound

This paper cites Perception Tokens Enhance Visual Reasoning in Multimodal Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.856611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.856611Z digest=sha256:eaa89e31008cff9bf6170075a11ad042dddbd72249d20081076ecded7df61afc

Observation 9ecec8f3-af2c-44e1-964e-f995accd4dbc · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.906301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.906301Z digest=sha256:cd746626af71359b0ae6cc5b5a4b0e540187797756b6a93a298eb1d4ee0ccc56

Observation 023d15c5-a23f-41eb-9564-d013bc58a441 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.930951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.930951Z digest=sha256:5bd3ed1365fc141417a263256ea6b028dec571769363280c02a2f817f87abc2c

Observation ebac17ce-1efc-4864-8848-c04bf13502bd · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.969152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.969152Z digest=sha256:a3ea081fb2dc4885ef41731e87161a4fc415cb5927659547046395ebd1486b9a

Observation c9ff8af9-fef5-462a-8545-4a856cbd90e8 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.014182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.014182Z digest=sha256:e4c7cf7f83c8b8cc0dc2292b8bd18334efce74a2303c15f2962d5a74fe21a8ba

Observation b0a05cc2-2c03-4da4-bf31-5f88bf5219d4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.061180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.061180Z digest=sha256:d21f632457b79a6cc467f90f1cb3c875b3da746881e59899182a6fe073631960

Observation 4d17a14d-d0f8-4327-9080-26d7492ad706 · outbound

This paper cites Gemini 2.5 pro preview model card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Gemini 2.5 pro preview model card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.108675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.108675Z digest=sha256:a3bba4944769eb862a3adff216042d5fd3541a0ac3292859a724ec95d002985a

Observation f1e0314c-df3b-4a52-8752-6ab3de98c9fb · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.149124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.149124Z digest=sha256:17390532e0694371dcbdd4b1387cd0f2a8a40312afa002f6adca248f612c46dd

Observation d879ea55-65b0-4411-9b04-86a11e4af277 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.213124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.213124Z digest=sha256:a872045a79f6ba9142a8247ac284350cb6eb3acbdd7c31166006c56d6266a9ad

Observation c25cc8e7-7529-47dc-9568-0ef4fd50a058 · outbound

This paper cites Cantor: Inspiring multimodal chain-of-thought of mllm.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Cantor: Inspiring multimodal chain-of-thought of mllm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.465045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:12.262102Z digest=sha256:6624f8574d05bd28bfedf99765f5fa2458fc35892d82ca359bed9489888bf997

Observation f24ee158-fd32-435c-8c58-fa50cdb828ab · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.301566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.301566Z digest=sha256:7b2c95fc42a9eb50afa8c43db799d1d2706610f98fb6e128b7647753531dc636

Observation 5021add6-3abe-4e19-9a91-c1f21a66fc32 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Measuring mathematical problem solving with the math dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.290620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:12.345211Z digest=sha256:a7995025f7cf5b20607016cc2470c53caf898fbeb3e405765b6d2e65477d90e7

Observation 7efaa70e-e358-41cc-8c91-60c873221cd6 · outbound

This paper cites The abduction of sherlock holmes: A dataset for visual abductive reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking The abduction of sherlock holmes: A dataset for visual abductive reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.168420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:12.398632Z digest=sha256:0b9c5fe8365124dc52e33f95ed0d2e4f5a2e9c58f3b6c06c13b332d9937b8796

Observation f35b0048-75c2-496e-b003-074ef2392571 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.450962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.450962Z digest=sha256:363199f6b659eb13c5549e0cf74b4a5a8d0fc212ef983de3d55d00b4fd48d485

Observation e41420bb-eee7-4219-a427-1d2970955726 · outbound

This paper cites GPT-4o System Card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.506068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.506068Z digest=sha256:82541f06273e390e3c9ef95a5ec24f49c3f2b95e2a42d47a5cfa2e0be40c5ff2

Observation 11301c7e-0e1d-476e-9355-b371fd77dd1b · outbound

This paper cites Math-Verify: Math Verification Library, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Math-Verify: Math Verification Library, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.006825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:12.549003Z digest=sha256:2d933615f2ae058b6a874bd638a633a1e5c7db650b733726a637140875c8eb36

Observation 7e1d0daf-db50-469e-be9e-1c0f31cb1850 · outbound

This paper cites OpenAI o1 System Card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.601005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.601005Z digest=sha256:b3737a13ad6d6a1e34d7d8b0483224e69755f696f4aae38376a433613c91ecb3

Observation 46c18466-ed5b-4404-b7fc-2097d1c65101 · outbound

This paper cites Abstract visual reasoning with tangram shapes.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Abstract visual reasoning with tangram shapes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.814552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:12.636382Z digest=sha256:f1c3b28e8544529813457d47a88def7dc0fed88454d3d9dae4217f83d303339b

Observation e0ac755b-9f00-462d-ab09-3b38b42ce4ec · outbound

This paper cites Dcot: Dual chain-of-thought prompting for large multimodal models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Dcot: Dual chain-of-thought prompting for large multimodal models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.642575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:12.689346Z digest=sha256:c82614ff2ab0e75bc523a2f9bc2249fb99a79635004d0ff40a413b619919f14e

Observation 3b2c1749-b529-46e6-b946-27af2bc38d7e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.734446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.734446Z digest=sha256:cfa6c995d470667938b25668d518abe0b0e6c3f46d0f019028564e0538702d83

Observation 0eb33f72-d27a-4a76-930a-3606de24e0fc · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Silkie: Preference Distillation for Large Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.792022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.792022Z digest=sha256:9d7ca420a8546daa1d0564fefcde9a7e3da6a968ad9cd26b436a1bdc9a008fa4

Observation 62f27d10-92d2-4000-9c11-92909d519264 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.887272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.887272Z digest=sha256:9c9ade438a9803040de75c5f4faa8d77eed3697c409e6e377118e9a31a12f71f

Observation 1208e7d4-dfad-4aa9-ba0d-592b73ae6df9 · outbound

This paper cites Diving into Self-Evolving Training for Multimodal Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Diving into Self-Evolving Training for Multimodal Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.938364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.938364Z digest=sha256:e3977cbbd2f84567eae11dff6963e921ec9aeb5dc414e36d15ff5c64c32a125a

Observation 3194040f-a74e-4226-818d-3234125cc836 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.978031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.978031Z digest=sha256:5f9c230d2914e2e3896aeae3f1d57b18191b9cc92bd3b679b77965257f0f3a50

Observation d0bb0f6b-b5c6-47ba-b795-aefcf70039dd · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.510804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:13.031238Z digest=sha256:f1086a6c88b33e677aa5e8ef1b7d5316aab9ec4d691b2df633501511261e13a4

Observation e43b2e78-e870-4fba-83c5-a9fc52b45e50 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.090908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.090908Z digest=sha256:ac1a8f3d1f5143aa7c1a43779cdf7c2dc599980c99824c839262b0e8d2e8000a

Observation baf9e354-97f7-452c-aeb2-9d34a4aef0c0 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.438773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:13.159210Z digest=sha256:2f873b0f1461bac0049ec9020617446879e6ca1f74105f2c6df93ac585fd6b5d

Observation f3ea50c4-9b66-4532-954b-befd96fe0118 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.202412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.202412Z digest=sha256:b7121704ea1cb16b80edfe532857d8a62dfd487ed7c018c7cea181711438980d

Observation 768f4ee9-819d-409a-82a3-6511bcca86bf · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.230798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.230798Z digest=sha256:94517efe147f3a4ad4cddc37f02038d9afdeff374aa5ce3d824669a4f9cade07

Observation 85d412e4-cd40-4c36-8b56-4669993426c0 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.261487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.261487Z digest=sha256:6d94fe837a95a2d9b40e5003d3f5d64819211a0e3fab97903e24ed99257b1200

Observation 86254fe8-b730-415d-851c-388c2eae1875 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Compositional chain-of-thought prompting for large multimodal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.282182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:13.311730Z digest=sha256:3ed0a6bbc7882f9a9e8ad4d554d903ca7047a6ee25df5ce4a604b95a157a1e7e

Observation d1e06d9c-a647-4703-aa0f-1712997adb57 · outbound

This paper cites O3 and o4-mini system card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking O3 and o4-mini system card

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.169036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:13.342588Z digest=sha256:0583a81d7d40a7af9da65c5958ad9342bfb2bcaf2d41ee6ec5726e4a9991915d

Observation d2cd426f-ae60-4c5b-a659-a910706b907c · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.380294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.380294Z digest=sha256:7dc5c98e7a54da6ba66aadfbd4648f41f48a5b378d18f40e0c62a1b2f35500ef

Observation 7f4e8530-e3d7-4994-a8a2-82b77e6730e9 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.434001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.434001Z digest=sha256:3ccbc9a7969e9795725c41cc412f8f10ce0b41570b3f2d99a4c522ef8cabd6f6

Observation 7b8c2e92-d821-40c1-b36f-71bedf13125f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.463903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.463903Z digest=sha256:fd62539cee7d8a1c03857b7704f19f1d9a4b6a9cb81c511a5020484125a04022

Observation c6d3d20f-5991-4e78-8cee-5c35904d0653 · outbound

This paper cites Rethinking Reflection in Pre-Training.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Rethinking Reflection in Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.515733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.515733Z digest=sha256:cd28b2ccd771a313d79f861a437e62e2f83e3a57737f4f1c413dd92607f91215

Observation 9ab4e716-36fc-488d-9445-0b2094067eea · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.537018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.537018Z digest=sha256:6200aec8b445fdc8b03264f8a9f93c318e5b9471abeabd4a1a61b6e156fda8f2

Observation c9f8922b-8b8f-40fb-ad7f-a5e4a598770b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.586726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.586726Z digest=sha256:90fdb1a9be7d8e73cc8b53519ea639421952a3eb46557aacb97022b89929085d

Observation 983461eb-e95b-4f81-bc9b-e7161ec4e2e9 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.667498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.667498Z digest=sha256:6914c2a66dab50d15cab497266341f42565a23f671083283dfdf3c7159ab573a

Observation 9623ce81-8064-42c0-b353-2e803c167ba6 · outbound

This paper cites Kimi-VL Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Kimi-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.706010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.706010Z digest=sha256:1c644fce3617cf208a0535f81f0383642ef147a1192bc2fc7aa0b3f7ad27a561

Observation 443fff95-c174-4f1b-b701-1aa212e2e4d1 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2025.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Qwq-32b: Embracing the power of reinforcement learning, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.751900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.751900Z digest=sha256:e39a19e1e240201c8183bd72aaeb74714b3eaabfb2af4451d8235f53b4f3f751

Observation cf62bded-4187-46b7-9d7e-cff9c6edfacc · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.783825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.783825Z digest=sha256:80712efcc1b25786a9d92a41d647893665aad5553c36b0e9aa3824a18693f7c0

Observation 9a19dd4a-05ab-423d-8799-5b2d39bfd03e · outbound

This paper cites Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.815176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.815176Z digest=sha256:5ecc7353a01419920beb3b15b418856bd7ce5a9d0249d4a17e964f9e4511f55a

Observation 7e2bbec1-3765-486d-be83-792e41dc3a9c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.847422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.847422Z digest=sha256:39bc557402f7f67a9d8a03530e92b6a0323a0ad0b237355507352fcdb2e18613

Observation db4734aa-bd8f-48fb-8b6d-ad871887ec71 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.922969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.922969Z digest=sha256:76c8fe77428f5963fca222ad8dc6c422b8ebfec84d96995f0638fce6e4eff4fe

Observation a17f981c-2931-4cd4-a188-00780aeb2972 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.976791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.976791Z digest=sha256:b5224eb362f749ca8164327fd474db97eaedbf817ddb6c3b4a0a62d08c22c54a

Observation af52cd2a-3648-4a19-8272-91c349c85350 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.016315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.016315Z digest=sha256:12c6e97a9b3bbef09754373929e767dfd50aaea2cdfa58141b88b503ca4cc540

Observation be5d73bb-39cd-4d6c-9ab7-4903ac947a3b · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking V?: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.047663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.047663Z digest=sha256:740a581fb539ff8e4c0a5891778de76111d0a70f6d89a03ad66f815c3adf3169

Observation 6f645a90-1413-4158-bc97-bd2620bea1e5 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.082910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.082910Z digest=sha256:21cb14ff60a3c9344a6c7b15b5381db087e71f50c6d032b8bb032b17b76391a2

Observation f7ffcbb5-699a-45e6-a7b4-69e78ef5f9e6 · outbound

This paper cites Valley2: Exploring Multimodal Models with Scalable Vision-Language Design.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.123082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.123082Z digest=sha256:36daccaa55d4aefe48afad448c468a4a942c7f153de9d9ea001b020d4c8a691d

Observation 325687c6-cd93-4575-9e89-36946bb2e03e · outbound

This paper cites Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.044283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:14.154557Z digest=sha256:cd0fc34ba3229a7c03446a4fdd151359099aeee8327afd8a2502330d735e86f6

Observation 8216b34c-7b24-4c7d-b699-180713ba760f · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.196592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.196592Z digest=sha256:f880ab2e1f5fc3a251e48f1c85757fa5227147f8e23d8c8a27afc2526af70192

Observation ef1a0837-1386-407b-829c-63f946f193cb · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.243341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.243341Z digest=sha256:3a88463ebdb3f5b968a82eb3d1f8cb6f4f23811a3b1b4b654a849df8a861e974

Observation 126c7a05-0e26-4e47-af71-db622ff377b7 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.280040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.280040Z digest=sha256:ab79a7773cb6e94c4cc027218f44b440194a53927ca11a3c9c41924b8e3c276d

Observation 5cb9a2fa-18d6-447e-a62d-cf036c3f1d21 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.320981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.320981Z digest=sha256:10f7921ada411f9bad1cfe815781748efbdabd9409131ae09a95e4d08651cd71

Observation f2a6f081-263a-452c-8fcf-e02f9268efb1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.363044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.363044Z digest=sha256:46b930fd93092d1a8ff93ec862253cfb6733db2fa7edd47de6a4b5fe780d564f

Observation be287d4e-fbca-4dcb-bccf-494b8b5b684d · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.405560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.405560Z digest=sha256:c272be7d0191d1622e474422916dcce43a9ae6502077757291134322a4ad1513

Observation eba62b12-8943-4eb0-90cc-1e83da13778c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.467663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.467663Z digest=sha256:5eb45226883a67b190e54811309a4431f4a7c374ad9224b59c5c7b3f4e97e855

Observation 5fc1b220-d50f-4c6e-8177-b57b88ed8c5e · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Griffon: Spelling out all object locations at any granularity with large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.944518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:14.511694Z digest=sha256:6097c55aa29f8a37ef9a986b59a0d221e01a1232661d2facdb75f019dc431e27

Observation 2d6a7a05-c0da-436e-bf2d-44ed9fd03936 · outbound

This paper cites Ferret-v2: An improved baseline for referring and grounding with large language models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ferret-v2: An improved baseline for referring and grounding with large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.766547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:14.559871Z digest=sha256:184799ba55f39950b3c1bbd280d82ce884744f1ffdb5dfd449f48cf0662ac340

Observation 9d312232-0a92-41c3-a7d8-a8b6e7113e02 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Improve Vision Language Model Chain-of-thought Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.605390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.605390Z digest=sha256:c83c4d03c844a296023b13c0f5ef52e50328988528c8e746ca881c1556d51bc3

Observation 4a5719e8-17ef-4d2f-994f-0806687c0102 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.660602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.660602Z digest=sha256:bd87dad9fed247b449d17b3a2490aee31e4711541b94bbbd68aa7516b8a75122

Observation af67a684-595c-4119-b025-89b1851ed955 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.702270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.702270Z digest=sha256:b7a9faaaf2f8c13e635745691fd8ab87e950ca6109b5e4948a6084e1162dcb47

Observation 5e94eb29-f3b2-481f-9344-cd4227812f63 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Multimodal Chain-of-Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.744601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.744601Z digest=sha256:38a57db097c1650a4a0e1b60c53194b6ef0244a3574d011e4d62fe218302589e

Observation b61ee54b-df04-4705-99f2-a932e7f313c4 · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.808518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.808518Z digest=sha256:9fcfed2cc4c391750df2cac07b0416d16f470cf1f68e214ccd7627f1ac748d0c

Observation ab1a015a-c659-4ae2-979b-d101bc05bf37 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.847482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.847482Z digest=sha256:7e8a04c9ccf1daf29d9ffa4d16d37aa3d85c399b2d99178d0143fd700778ee0d

Observation 154118fb-2756-4d03-a647-3194e61ddeed · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.641879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:14.880283Z digest=sha256:df725d411f92c6a2f1341a0874b3f88e45bd918c86d26a8eb37836c2addb84cd

Observation 49b668c6-1839-40b1-b341-9a14a4da1826 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.545042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:14.915753Z digest=sha256:996fe02a0d1111447b430031be6990421c7ae0bba12168cd3b83cfe998842020

Observation 8b139858-c0d0-47db-92a4-2475a2512a97 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.460148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:14.970907Z digest=sha256:6374d86cc190b2db270c0fe036f30d571b7fffc72a9320d72b113855ff059eb0

Observation 3bb65ca0-0f86-4c04-9581-9da1db0c5bde · outbound

This paper cites model’s chain-of-thought (CoT).

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking model’s chain-of-thought (CoT)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.375472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.015345Z digest=sha256:47b39066438ed1597b9b8af7a596aa51cf9d754119c0fd77eaafaaf1f89cfaa2

Observation 0337b9b7-4972-4951-80b1-eb404c4cef5c · outbound

This paper cites - Then, wrap the model’s entire thought process in <think></think>.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - Then, wrap the model’s entire thought process in <think></think>

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.288363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.075607Z digest=sha256:648aa1a45bbd7034e4c30dd5d8b6df8cefdcf1f2d32122ebe596ac94a166a463

Observation 03b5c2e2-ab2a-49f7-85a9-c3c61f023780 · outbound

This paper cites <vcues_1>, <vcues_2>,.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking <vcues_1>, <vcues_2>,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.211611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.125236Z digest=sha256:dd7c01b9a3aa75eb5b06697275c935e038079c9a6cec7c41be93dee5ca5bcafe

Observation 92a74542-590d-4e50-a054-f3ad6876cce2 · outbound

This paper cites All the data to be processed now concern reasoning errors based on visual cues rather than errors in visual cue perception.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking All the data to be processed now concern reasoning errors based on visual cues rather than errors in visual cue perception

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.139294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.156197Z digest=sha256:9f4ee1f3c851faa14a49102d6f5665571b437c8a9fc6bc1d2584c32c84319de8

Observation b2715fd1-294c-463c-987c-a6d13ec84414 · outbound

This paper cites based on the rationale.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking based on the rationale

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.067581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.182656Z digest=sha256:89195a6a90a642f68f97605a7b2d273d9cb8955a999ec21d288200e66128fbf7

Observation 9abc4366-1090-4d11-b7c8-1a1192610bac · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.981376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.223161Z digest=sha256:840a90f4ff22c5d72357dd246731995dc6ec41e6e67706a4d9bd4317a66d91f5

Observation b3501fb9-f946-451a-ae49-8aa9c067f709 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.906529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.258999Z digest=sha256:07e70bfbf049073dd38bd5c28ee29c74709bd055aaff49020e92a0b1de70aedf

Observation 0ec0a9da-6c2b-42c2-82e7-bd83727448e0 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.830330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.298530Z digest=sha256:677de865f17e17f71507a849691161d7b839ddc2b1950492a275cd65d408309e

Observation 7efba94b-a213-4ed4-a25a-0e79e112b29f · outbound

This paper cites Let's verify each visual cue and its reasoning before finalizing the answer.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Let's verify each visual cue and its reasoning before finalizing the answer

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.740232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.335118Z digest=sha256:a393e61a0eb2a2749678526c896a614de6e176b182a13bacea3e926fdf3e8010

Observation 37267687-4afd-40a6-98bb-e09353490526 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.651087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.400221Z digest=sha256:75321c4d5dcaecfd37f4e3d6fe32ffa438bedf99cbf4250f96713ea36a63b187

Observation 8c3afcc1-d706-4bba-8d3f-c0c4259cfd39 · outbound

This paper cites - The angle 78° is an interior angle of the triangle, and angle 1 is 42°.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - The angle 78° is an interior angle of the triangle, and angle 1 is 42°

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.564150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.434526Z digest=sha256:eedc5046d4a5ca8cb4ef330baf4c6161d3af96b5923adcac56516a94d69a5a3b

Observation 1993d855-5ca4-4b5a-a7dc-9b2a6750ef4a · outbound

This paper cites However, upon reevaluating the problem, it appears there might be a misunderstanding in the interpretation of the angles.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking However, upon reevaluating the problem, it appears there might be a misunderstanding in the interpretation of the angles

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.470842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.463161Z digest=sha256:ba5a5345cbf0bfd0b2801567470b29fe9ab310c0ff7b90c33eaf9daec28e4b2b

Observation 27496727-abc3-4c17-824f-ace683444b01 · outbound

This paper cites - <vcues_2>Angle 1 is 42°</vcues_2>.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - <vcues_2>Angle 1 is 42°</vcues_2>

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.279875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:55:15.498636Z digest=sha256:a857355c8a8b0fa20eded8f79be3c135797b079a79e85b21d7b0a0d684e80ef8

Pith citing papers

Observation 9bf62d37-bacf-46c9-9528-1e292c4f2fe2 · inbound

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law cites this paper.

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.575834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.575834Z digest=sha256:060c53a852cd7a6e87b26cd70fb7711081e592842e50205ad2c824b8d63f1435

Observation 51434a38-27d4-4cdd-9a2e-fd34b8e95e25 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.600332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:2bf04bc99407e11a5650167dd370997a3a2cc5e407137928d0d3234b582875f6

Observation 74e75a2a-7660-4070-9f2a-35d038cf8ac1 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:46.730039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:46.730039Z digest=sha256:826d41557e553173d1cb71dc242a8911b233b0d94a39e5871feacbb869a3adcd

Observation 3b5c11b0-909c-4144-9dbe-7e7e7e6e00a9 · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.576402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:93a5a1a4376a116dc6d54aaeb5c27db046f9eaae517d56e9b71f7a8ae8e17b74

Observation 8160789e-f389-48e2-ab36-3b6de0efeffc · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.576609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:1085226c20d48dce1584d86926b6399f992de4f19aa3d2ba7f8a272a2f0dbf84

Observation 4ea4b481-5c01-42f1-8ad5-12453ff09425 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.333992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:9fb4627a70348a3aa1b106b1f701babe37933adc820c88a640b98bbdfcc55a82