Pith. sign in

Paper Citation Record · LEDGER

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

As of 18 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 4 inbound Pith citation observations for arXiv:2508.12680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12680 v1

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:24:08.529010Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:09:33.968186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:06:10.238711Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfd5f7f5-cedb-4235-b24c-3b2b79e6a404 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.025253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.025253Z digest=sha256:d48cb3f67334eda796087161f540b5341aeb628d1a3f467fdc1b4773f6930bb6

Observation ce8eff9e-56fc-4499-ad9d-9ffb88505115 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.031605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.031605Z digest=sha256:a7fb64ee1741996f321b576da11bea38efcacd9411f4aaaadf95efc680507250

Observation 6aa7cf1a-2207-493e-b21f-e67f780a5272 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Measuring mathematical problem solving with the math dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.037031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.037031Z digest=sha256:2ba0af5587c07051c582801fcfe8508cf4e433e66e08d1cab5139659e11666d2

Observation 4a2cf671-3e03-4208-86e5-e1bf28c41528 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.042129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.042129Z digest=sha256:5601838ca2435b3af3688c1510bc4ecf5a3741b36178f68890b06f54f41b9be0

Observation 5aea5da8-1175-4b5c-a206-e8dae25642b7 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.047280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.047280Z digest=sha256:5abab2f4cc988d56f7bc4433d96de8379e8edeba166efad136be9176cc9d3fc8

Observation c00739b9-2aff-4843-bf97-9b9bc98461f3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.052248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.052248Z digest=sha256:160691384128438af67b01e641fd34ad26d56f2102fb8687b35fd6d0264c0cb7

Observation 464d6984-42b7-43e7-980b-ed5dd2d088d2 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.057500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.057500Z digest=sha256:138a6a8ab03a65097d821b98a6d89b5ce82df931dbff4854abd3fb538b359839

Observation ade186c9-82a5-49a8-be9f-a10ccce1f593 · outbound

This paper cites Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.063540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.063540Z digest=sha256:5e728b7abec5dfeb8922b4349528a1ae53e28c2351e84f68591bd7bc30ead8a2

Observation fe4eab4d-e348-4dd7-85bd-740000ab12f2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Qwen2.5-VL Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.069223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.069223Z digest=sha256:01162cad6f4a091aecf55e85a904f80be012dc547e90530a858e632db20a8789

Observation ff11dbdd-89e4-4039-b08e-30fd1a15485f · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.075537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.075537Z digest=sha256:4f6c752d498258975ad123a1fd4cfeb99709c8e4c40c3ce7878ee6759657c449

Observation b335f2bc-d348-4448-9e50-7861a3efc2c7 · outbound

This paper cites Visual instruction tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Visual instruction tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.080483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.080483Z digest=sha256:ee0825efd14674062837134f07d2d14c7e6d4d7771dd1c095f0a19d3257caf57

Observation 58da3d48-c780-4ecd-ab43-50f69b2a6eb4 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.085320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.085320Z digest=sha256:90a5de122942f3e44c8b4fab3c1dcda6d2e5472abe5240fea8fa0800217f5983

Observation bbcb08f5-158d-40db-98bc-4747335ee78a · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.090540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.090540Z digest=sha256:bbed4be537fb37093288ecdf80c7fb0e1402d07ceb13c4a8584e95a81f62bb06

Observation 3bde4210-ac89-43db-a1de-bf6ee3940371 · outbound

This paper cites Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.095874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.095874Z digest=sha256:f8ebf17dc3d82d3df8d87954ac8d4af6c0a8c50d98be2ba636363ba87913e210

Observation 7622440a-edb0-41fd-b314-135e30cbaa0c · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.105506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.105506Z digest=sha256:87304bea91dac18ba0d180db6d011f87e5e32e35387e8e1709cad5b83cbbe8c3

Observation 112b9614-e7a1-4a83-8c93-a72286806a15 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.110787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.110787Z digest=sha256:43244c83b59a9d34605d949ea4f9004731d6d82bc5c38fbeb245939ce8812d67

Observation e23391ff-c8d4-463f-a79c-78f2522cadd0 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.116175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.116175Z digest=sha256:3b94b0f45c43bef01dc6853f370cd13e255966ce80e735a24d8dc7451ffb3478

Observation 6e8cb2c6-8893-4d19-ae6a-e7dba4b1234a · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.121024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.121024Z digest=sha256:f80878450e23f017a5ae6911d2d57e0f65b777963ff0eff4dc420844d190598c

Observation 642c9b11-9ea2-4095-8624-faaa2e2aafde · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.125709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.125709Z digest=sha256:93ce9f27634b7d63a76b0fce3706bb8d2de5a649eb74f1c732865359a48b1c7e

Observation 90239c86-c51a-4bc9-a116-367c475a62c1 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.130682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.130682Z digest=sha256:52e794a8d1d78f8a6cc125e898af49d0c37469bbac069f842b0a89d1c3ec7844

Observation 0544f2a6-a603-43a8-804b-37d51e1ee67e · outbound

This paper cites Kimi-VL Technical Report.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Kimi-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.135546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.135546Z digest=sha256:7c080e7fc47eb654bd805281333bdc4ffa045b55a01a64dcbad09d9d5352574b

Observation 2af91df2-18d5-4b30-88f3-717e80af052b · outbound

This paper cites Improve vision language model chain-of-thought reasoning, 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Improve vision language model chain-of-thought reasoning, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.140526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.140526Z digest=sha256:5eaf1b35cba9a93bc0b3a8c263915f9071c20487c4da96cdc32bd003b5299edc

Observation 643b81b0-db47-446e-84f0-a9d29287b31c · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Llava-cot: Let vision language models reason step-by-step, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.145289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.145289Z digest=sha256:99f49f888259c302ca5dda1dc8c3261f3c6ef75668f8e40de47ad9e181ece9c4

Observation d2394a16-e466-42f0-a81e-1ad9d082294d · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.150221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.150221Z digest=sha256:10bc546b254b8730067e4edf851a3a0eb9c385cd0ae6a5fc278824453ed2c8af

Observation d2cec0f4-2c7c-45ac-8c25-2662a84eb841 · outbound

This paper cites Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.154904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.154904Z digest=sha256:b0bd3afe6e06a410c49f4442d5ef7f7032746a690e2927c6539848f009b3e67d

Observation da7e05ac-18da-4499-9480-2d1f810a8482 · outbound

This paper cites MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.160749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.160749Z digest=sha256:febdeb463b23d029a53c9176e3913624edfc0a6b49b0c8c3f62433791bfd27a6

Observation 17e8a213-d68d-4a21-b5bd-71b06b1fe91a · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.165737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.165737Z digest=sha256:cf12d8088df5e83f113c9f9d37ba2066767212e383eafeef32ae8d5cc4c542e3

Observation 183598ef-8a3e-482c-8f50-d146f36bd236 · outbound

This paper cites Estimating Training Data Influence by Tracing Gradient Descent.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Estimating Training Data Influence by Tracing Gradient Descent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.170718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.170718Z digest=sha256:772b219e9e896d15280e9e0575c86c18f7cf15ec017351c07c4c1be61548a0e2

Observation fb1a6e3f-3679-4b9f-a601-dedf248447a8 · outbound

This paper cites GPT-4o System Card.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.175673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.175673Z digest=sha256:db6b8be356a73fe088c2a411f51b698dbc6a9192d7291b7c0ff7e41d9d1abae4

Observation 606f5ab8-6bd8-415b-aa7a-9850ede75773 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.180238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.180238Z digest=sha256:02883e033e850bc74846f4238f850d45c2252069349cce2343752ebf220b7b3f

Observation 6251beb0-6008-4404-8d7a-7d8c47d1a376 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.185246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.185246Z digest=sha256:8490a01c55f050d0fe032f03d32ab379756325281b3b2d54d2aec36dba1870fd

Observation a7eb8a0b-c7b5-4ad2-950d-304f8a234132 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Qvq: To see the world with wisdom, December 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.190465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.190465Z digest=sha256:53de442d9bec4389931809c3aa6f502fdb8360e9a290407472e705b0dd867ac2

Observation 68b71451-4a7a-4856-86ba-b191f7adc41a · outbound

This paper cites OpenAI o1 System Card.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation OpenAI o1 System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.195348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.195348Z digest=sha256:03a8f910fe6745fc84985128da8e770905d3d4ee8802c43018dbfd3a657995f8

Observation f24ef1ec-9434-4f03-a185-96858763d27b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.200230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.200230Z digest=sha256:04ebb14484c5afaf3802c356c742867ce11f17921f95f8c22381ec02c262d135

Observation 88f412a1-b5a5-45f7-a2bf-78a6b741fa94 · outbound

This paper cites an unresolved cited work.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.205225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.205225Z digest=sha256:62d5eee18e4706c8fcfe4886ec7b1be7dd2c4c8de6824464a0b582895c49ac01

Observation 5d4b8d6e-260a-49aa-9a78-6e1471136246 · outbound

This paper cites R1-v: Reinforcing super gen- eralization ability in vision-language models with less than $3.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation R1-v: Reinforcing super gen- eralization ability in vision-language models with less than $3

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.210995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.210995Z digest=sha256:33a492b30a01b81084fe9aa1db0154ddd731729fb8fb25fdea12720db5a8e93d

Observation a3c2899c-091c-4675-aa56-5eb2bec22a5b · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.215793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.215793Z digest=sha256:5cd4f369a77002cc392608a7c909a45ecd8660200a5e2766cc4ef12ba6c437cd

Observation fbd8b0af-21e1-4a8c-9841-09f1b4e70411 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.220934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.220934Z digest=sha256:6a3b70ba6d533f74d0843cab4a81430a319cff0c9e847b75f98f81a3cd86df9b

Observation cc2396e8-cd50-45ba-8cd9-3706eac66bb3 · outbound

This paper cites Lima: Less is more for alignment.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lima: Less is more for alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.225981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.225981Z digest=sha256:cc3ff2ea3c896a20c5bdfef64910f2cb147e5bc7383656dc774e61d3a5421bc9

Observation f35c73f3-39d1-45fb-807c-6ba3f365470f · outbound

This paper cites Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.231109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.231109Z digest=sha256:24d43aaa8029bb85aafb91c0fda6ff3f350f7e1476e960d3bf773594896eee21

Observation 12946d17-c9d3-47d6-b083-e47acee5baef · outbound

This paper cites LLM-Assisted Code Cleaning For Training Accurate Code Generators.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LLM-Assisted Code Cleaning For Training Accurate Code Generators

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.236832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.236832Z digest=sha256:a37992a83e508d9255a92651bafa24e6e923b631f9bd9b962ff02dd2cfa06367

Observation c6e8cc79-7e3c-4088-a071-b849fad1b69e · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.242672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.242672Z digest=sha256:6f92097a894f2978e101d0888c51e37dc746bde4492757f78163e1fee99a7298

Observation c73c9da2-a9c0-46e3-a726-cbb2d4663d9a · outbound

This paper cites Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.247730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.247730Z digest=sha256:e9688652836668ba9a963cc0fd3569ca64be59973b1d8d1fc03bef28756dc147

Observation 0cb1384e-8139-4298-8508-cca5bcef0348 · outbound

This paper cites OctoPack: Instruction Tuning Code Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation OctoPack: Instruction Tuning Code Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.252677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.252677Z digest=sha256:b942545c8dcbc8afd5472ebea922d22d6440e8fe0d942dc0c34ff94258ab00c1

Observation 30a91729-8420-44bf-8e9b-39f6759bcbb0 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lora: Low-rank adaptation of large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.257664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.257664Z digest=sha256:563bb4f45842b29f5721c911b9c901e83795085224c919440bdff4e8a106c595

Observation acb0228c-59ed-4336-bbd5-7f6fc4c76ada · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.262872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.262872Z digest=sha256:8383e10862032ba7ef9ec4c7a2223dd750f55c4f784c28dc7545bcdd907ae6b6

Observation d4d2e983-fcd9-4d5b-808c-e86268208079 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.267817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.267817Z digest=sha256:c2c7f81d5fd39cf217706455812694e20e16fc886c20d19c1bdb6c891524a391

Observation 46cf7362-e5d1-46a6-bb99-4f1ba45e377c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.272810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.272810Z digest=sha256:b863c2ca6783363aa0547d7c0cc2df383c151deace05ede26d69c24025b2fbeb

Observation c5f39aea-717f-4e37-9215-6e84e5962235 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.277999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.277999Z digest=sha256:edf7b785a6183575e37ac2a1ad2ae486324af11360ad0c84008bde2c48f6cf03

Observation 75c08491-55e3-47da-af83-9fb8f069cbc3 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.283136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.283136Z digest=sha256:79350bb487268b154f3fa8688a5b506f86216e5394ce23563ba61ef548339ab1

Observation b9656623-4767-4186-94f5-b2ae1fea1b86 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.288087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.288087Z digest=sha256:d9d0cc9e7073dcef731a4c076bb360c60f2ae099da36c87324380a2f7e434c5f

Observation 697d2f21-c6f5-45bb-b7ca-7fd28bfbefa6 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Measuring multimodal mathematical reasoning with math-vision dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.292710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.292710Z digest=sha256:19d88e2eb669f5d81c844eec6a33c49df3860f6f29657bd14fc09b7bd404dc62

Observation 065bef9b-408c-473e-ac24-3904ed4f939e · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.297641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.297641Z digest=sha256:6beb9d95fddf03fd697876d098b442397963d0c9771fcbddd432788f2d1def5c

Observation 16a605b8-022a-4945-9fd4-0bc8dbb33625 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.302455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.302455Z digest=sha256:5707d15b9ba942ee6b4205adb89cd4f07631c5db133d4d8b0ff3e748a2e72b15

Observation 659a2e93-4df8-4d7d-9d13-5bcad45354f1 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.307694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.307694Z digest=sha256:d35eda093a93f1506b038657ce8e395e7e6827691e448a56cba35609f1ff1d0b

Observation ab350932-1d83-4672-a1b2-6435f145c084 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.313024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.313024Z digest=sha256:52bd16d7999ea2ab1da08bddb81a9c0427039635991c2fd6eacc09947ba471f4

Observation 2a673b41-5ccf-4fb4-8a99-b9d678df3c72 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.318268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.318268Z digest=sha256:fab099d5862895a791a9256a79a99167e88d88968a494a364aff0a24272c238f

Observation bd9394ee-27c1-4d83-ad73-20f2d6ae886d · outbound

This paper cites ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.322949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.322949Z digest=sha256:a50c7f1daa52a012993de3892e0cfdc9ed2d973ddd09d02159b001a43d2cfb66

Observation 72bfa829-d95b-4eb4-abd1-7d58dce836d3 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A dataset of clinically generated visual questions and answers about radiology images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.327988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.327988Z digest=sha256:39ca662562af1294d5445fb9345f2db9a418501a907c579e1c7ca53c88369e79

Observation f841676b-ba39-40c1-b25e-5ffa31b25379 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.332746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.332746Z digest=sha256:397d357e04fbdb4dc676fdd2dd78b89e959c25bcaf3029fd1381f7d538bba981

Observation 4a3f783a-95aa-4417-97e9-298d15cbc6f2 · outbound

This paper cites Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.337699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.337699Z digest=sha256:a55c968c0cb066a43e0e843fb348dbdc938bcce9d7561c2ced817ac3dfb818ca

Observation ff7cb12c-657c-4e22-8879-62ac8b89aca1 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.342648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.342648Z digest=sha256:22d3c348c7a79726bf14cc18f7a6166b63b7efc85ca1635178a7cd927eefc2cb

Observation 72f64ce5-8e37-4f3a-b84e-f3d780d1ae00 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.347677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.347677Z digest=sha256:0b2c11a3499331d728913723d080458d75131b29c08c6300463afd88812fc008

Observation da29709f-a27e-4bc2-b702-7e3df9903af3 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.352870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.352870Z digest=sha256:662b00b5aa8403cb40eee69f8d762a515686a4066d04702480e222c0a326eb51

Observation 34ae2762-e1ed-41ae-8ea5-71819f1e94ae · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LLaVA-OneVision: Easy Visual Task Transfer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.357697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.357697Z digest=sha256:695917d9539f01b4b6b7e63a757d2b8215cef3ab9e264faa59cad1b42fd9238f

Observation f10e8a88-ecf0-443e-9358-9d7289054b12 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.362835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.362835Z digest=sha256:5f9b23823dafb8b04e11700f7212aa785ade0c6ad076c5436d37259dc39e753c

Observation 65c6103f-3e1b-40b1-8d0f-fd6070afde1c · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.368349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.368349Z digest=sha256:7d89fecb9ab0783955f479330e42049f97a5970da208154ea23525b7620ccbbf

Observation 0e724969-432f-4218-8cb3-d931bb78a8dd · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.373195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.373195Z digest=sha256:b38648aa267f748c5938ae9e5dee66d54cbcb0a7c773bb3048e346cef2bc84fb

Observation c6bf0777-e8da-4546-96d7-41539bd2234a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Proximal Policy Optimization Algorithms

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.378185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.378185Z digest=sha256:e2607e6d448e3095eaeec6f38c5238f4cc858d89b4172de3625bdb6caf03098c

Observation 376bacae-793c-42dd-9047-57ab352cf7a8 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.382930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.382930Z digest=sha256:10515f9e4493d18a95571b5a5d24f81cbff66cda73774f5c4e4cbd8a8c66a74b

Observation 63c1af2b-ad80-420f-97f4-2035f9b86daa · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dvqa: Understanding data visualizations via question answering

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.387919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.387919Z digest=sha256:46e99b275fb878595819a4ed19416b06394c0cb5505ecc43a5ac21fa5234d14a

Observation 4ed0157a-9167-4bc7-b5e0-d26bd703336f · outbound

This paper cites Plotqa: Reasoning over scientific plots.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Plotqa: Reasoning over scientific plots

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.392660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.392660Z digest=sha256:2838d230871a35e7ec91a058a75ccdd04ec58e31ed3fcaf03872ea4b63caf76d

Observation 96109899-f9ff-431b-b58f-a0ed11b218ed · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.397621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.397621Z digest=sha256:4e7ef94d13c8341a36c9158d1107489a29c3c042c5e3ffd4013d3bbc793a81f2

Observation 48732e57-7c6b-425b-8266-7a93b8ea40f1 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.402852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.402852Z digest=sha256:b7b9aa916337c82023a507ecce8e8e1df512bbf5b38d9dd7e96a39aa5880c1df

Observation 2ae55cf7-6324-4e57-b2c2-2e1e7fa2bfe2 · outbound

This paper cites ChartBench: A Benchmark for Complex Visual Reasoning in Charts.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation ChartBench: A Benchmark for Complex Visual Reasoning in Charts

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.407822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.407822Z digest=sha256:2cb517daa1d3e1f9caf07dfccbf6233669d35ba3030e638a34b372d51a067e92

Observation 23c48653-37e2-4852-b360-82b30dbae3e2 · outbound

This paper cites Unichart: A universal vision-language pretrained model for chart comprehension and reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Unichart: A universal vision-language pretrained model for chart comprehension and reasoning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.412542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.412542Z digest=sha256:a6fcc8c3c5ec59476fe21106b0149509367131f4f01363dbd30efe9b2f8e4221

Observation 6132d4de-611d-43a2-9dc3-0ffc6e463a72 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Docvqa: A dataset for vqa on document images

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.417054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.417054Z digest=sha256:c74529e51c65084117a2ea008767f5eaa3440d97384a86d07af5061a4c793e84

Observation 30ce4f23-7c12-45e3-b24e-577d02c4cd73 · outbound

This paper cites Harnessing Webpage UIs for Text-Rich Visual Understanding.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Harnessing Webpage UIs for Text-Rich Visual Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.422038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.422038Z digest=sha256:6faa5f747a8939eb227514aea39f5a3320bb82ca25afae647221c3c183a9b0d8

Observation 9418bd85-c530-41be-aeeb-1d4b5681b22b · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.427313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.427313Z digest=sha256:cf1ff053643f643b06f944a65229f18861ae4a8d5e8cc42edf78ee9c85fa2a09

Observation 769eb7e2-f057-4600-a319-7c77768b5187 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.432638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.432638Z digest=sha256:bd4203fb53e68c43a2b9c9529b1c5c4a096ec2d3ac3e865526231ca5e33834b1

Observation d08d20f2-47c6-4c1f-a5ad-588660b69e8b · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.437400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.437400Z digest=sha256:04b028e9065a83ad922d0eb84e9e10cd2ff0a717d1efad31135602b9393b1b31

Observation 3b5cebb5-9ff0-4f8d-88d3-4b0a296a8cf4 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.442245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.442245Z digest=sha256:e387a54ec5c9bd9a109b3db42bae6964504b95eb1ee6a5fd06280feaf0ed8457

Observation 84cebb8a-e84c-4ded-942a-a9e0dab82fb6 · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Solving geometry problems: Combining text and diagram interpretation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.447210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.447210Z digest=sha256:12113c80e074a6759514b965d5bc158c559afa9ef45dc85b4a27afe1c9083675

Observation 69d8adda-2460-48f0-be5b-5764e70775e3 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.452090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.452090Z digest=sha256:2d88b0a4f5c8b0d2ef1eafa919dc051c7fed0f414cd587d888d178896b3545aa

Observation 1ae9e6f8-e958-4a16-bd67-35a411a8f7fa · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.457298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.457298Z digest=sha256:09068e4fcc3990e8af820a984b9798c2a12340bfdc5160c031dfb35090db372a

Observation 3229e0cd-4c93-42b0-a050-a02eba852c68 · outbound

This paper cites A Corpus for Reasoning About Natural Language Grounded in Photographs.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A Corpus for Reasoning About Natural Language Grounded in Photographs

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.462326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.462326Z digest=sha256:b00a30bfb5e971663761f7c173c919d745b741fc9607d538a9166525c0bfeab8

Observation 4e6016ae-9dd9-47a4-b9a8-4917f9ef2ff6 · outbound

This paper cites Image Retrieval from Contextual Descriptions.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Image Retrieval from Contextual Descriptions

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:24:08.770841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:24:08.467219Z digest=sha256:9f94d66e5bac08ad08665c747fa5c83bb05fd5af9490497d1201df8ea9662deb

Observation 6a872cae-9d7c-4986-9b99-f5f33a6cec46 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lawrence Zitnick, and Devi Parikh

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.472416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.472416Z digest=sha256:a764730db03b4eff4c18627392d348669d29404e9e0acbef8b58dd0c10a19bab

Observation eb282081-0b1c-42d3-8e7d-198b15c2a976 · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:24:10.024942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:24:08.477274Z digest=sha256:e1e7f3ace1aa6f761669e7f4e0f99823d896f5a83c93aa35868d761862573bb4

Observation de48a35c-7bd1-4689-8db8-a768aece6e2e · outbound

This paper cites A diagram is worth a dozen images.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A diagram is worth a dozen images

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.481966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.481966Z digest=sha256:22b825ade53f0331f0d3e5baf75ae16e63aab6f1c3ecfe3ec78b6ea611b9b2b3

Observation d15722cf-d116-4081-a0d7-a43e7e8b4b34 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:24:09.996101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:24:08.486479Z digest=sha256:8124c24a58ba5bc66d00835197b8bf8d2b2d8ebec29929a9553dfc4a05e15d8e

Observation cd6f8cb2-0f48-4a38-a74d-5288742c67d2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.490975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.490975Z digest=sha256:08853b0cd93a31710967dde2064cc12e7466fd416fc046314bf8c4b97031ca71

Observation ab713047-994d-4443-8b21-b87b0005d038 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vizwiz grand challenge: Answering visual questions from blind people

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.495486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.495486Z digest=sha256:5c86b51c127697b72a9a3d7e2d5cd51b332946a4e3495f34d39381df9b2c9dde

Observation bfe0133e-91c3-4c81-85a7-0729729d3aa8 · outbound

This paper cites Towards vqa models that can read.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Towards vqa models that can read

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.500513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.500513Z digest=sha256:fd3ddca9eb14531f5aaee795e9f4772d59a2910427bc2a47c46339fda2d2d8e3

Observation b17bf1cf-2bfd-4532-bedf-bc28197e34dd · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A-okvqa: A benchmark for visual question answering using world knowledge

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.505250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.505250Z digest=sha256:ded4187c1be34c46438345ec3d8a29198be3bffdffce1ebd23ea64bc06bb3596

Observation e24b1632-9004-4fb0-b05f-7da7199d71b8 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.510153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.510153Z digest=sha256:f62b74c1a631371a39f63aae0ea97b6381b7b5056aad2a98bd459be64bdd0ce0

Observation 15866e9d-4e44-40d8-9ab7-a6b5d64790b7 · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.514831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.514831Z digest=sha256:ebb57ee088fed53dec834d2432a4418a19df52b1071c530ca913f3ff96f30d6a

Observation 0d26467e-9283-4c20-9eb0-d4c876cad2fb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.519580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.519580Z digest=sha256:3491753dcc102e107eb8bfb55e159d621944a2da6553493222b71b5c67ef6d8e

Observation f701a439-9b26-42c4-9d6f-f0b003162a47 · outbound

This paper cites Chart-r1: Chain-of-thought supervision and reinforcement for advanced chart reasoner.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Chart-r1: Chain-of-thought supervision and reinforcement for advanced chart reasoner

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.524451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.524451Z digest=sha256:1b185193a9afdbd64034c389950c5d8e8c501c8f4c2e81426e9d16bcc58857ec

Observation 76ea1cc2-7929-4419-bcf2-0c36ff67336e · outbound

This paper cites You are a QUESTION-TYPE classifier (do **NOT** answer the question itself).

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation You are a QUESTION-TYPE classifier (do **NOT** answer the question itself)

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:24:09.922336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:24:08.529010Z digest=sha256:ab4a615c3751e70e789c4214781af2f74844e05b8c41a103ce33bc27063e39be

Pith citing papers

Observation 6bf98c01-7ddf-4da9-888a-95077f7dad34 · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.744656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:dd1e122d5f341892b1fc1d910cb90d555d2efc13021be58153667361730dbdc1

Observation 27a7c529-19a8-4113-909d-20fb608e0500 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.296681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:60375e22d78963ce77ac1dfad17ddd25896ce0b2699a78232fd092c2434d41d4

Observation f3d80314-5b0f-4b0b-bdf7-8e8705610c62 · inbound

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution cites this paper.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.242064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T12:40:34.424341Z digest=sha256:8171fba3318d1e8f02c40df6662ee270c47483aa7280821781eaa5e331676502

Observation f4626655-0973-4564-9de3-00cec701f71f · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:09.398205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:8efd12627708f92d7f96dbb0b4edca4a62874cd8426766a3c2d64565317918e0