Pith. sign in

Paper Citation Record · LEDGER

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2505.24120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24120 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:03.076855Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:03.671614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T05:55:31.163509Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7584910-2c37-4e9a-b185-661879891cde · outbound

This paper cites Gpt-4 technical report, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4 technical report, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.070452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.070452Z digest=sha256:75841334fcea01ba30aef10d907096326685e7d1dc70cce3a36561c20c857288

Observation c7f311c7-8fad-4b1b-8402-395389ae472a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.813861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.118262Z digest=sha256:d4f8c5ce1c586a89934cc063dd88b59ba23c284c2bf4553f2d6a5c40c3eda790

Observation 89ecddcb-d994-480e-8a5b-2eabb2e6abc4 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.703189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.202715Z digest=sha256:0f5c5cd77de67043824ed2e783847b66b162410c97bd7582b44dda73f2e2158e

Observation 94fc0825-f2b2-4750-b2e2-1a5f9fdb828f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.273403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.273403Z digest=sha256:dd6a1646f3ca1fcf9690bd34a5b19dc1d621e4d2b18656626705c95fa54de746

Observation 30fdd865-079d-4720-a636-45aff8245f09 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.348853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.348853Z digest=sha256:ee216d6ecb899e80969bec6584b9e5dff367ead0858dc9ed6a95396ba1838fb2

Observation c7d9aefb-fa90-4664-b2b6-5d030929f230 · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Kimi k1.5: Scaling reinforcement learning with llms, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.582667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.418298Z digest=sha256:14ff6e731351e67c49662dcc5ed9fb59e62897dfe82ed11c52a55ebc10422f58

Observation 161e185d-2897-4d8b-a607-9c89c2732e88 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gemini: A family of highly capable multimodal models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.441409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.496657Z digest=sha256:b20f5e30a036d6f72ffb34a3fef27fc4e848890a266f8c6aff8b5e57b13ff1e1

Observation 324b1f69-6911-4cc9-b3fa-04e331ba0a3d · outbound

This paper cites Hello gpt-4o, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Hello gpt-4o, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.301657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.588916Z digest=sha256:3376fc69b8e1ec1048b52c299d95c648a4a00fcfa99e0e31947c43796b1011c2

Observation 4b54aee7-b2fb-4eac-b9a5-a85a64adee3d · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.154599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.657293Z digest=sha256:15f9544d3a6753041b55409a1fd9e002a7ac02192b8628587214c0df1891a8e6

Observation 3b68a956-1007-41cf-af05-323095274103 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:09.028912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.737185Z digest=sha256:78e28af9f11034ddbcf21f529b9bd73ec78bf3b4b2bd8db741ab6832abf994f8

Observation 691bd1e8-04a8-46f7-b709-c11d53c7ca57 · outbound

This paper cites V Jawahar.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs V Jawahar

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.888367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:40:59.885061Z digest=sha256:578b7435626012bfa47b5fe32888591ddc5c9bbd2399cb78e05c33aa172c3bdc

Observation 5ab11a5d-6aa8-4048-8acc-13bb2133af01 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.731263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.011399Z digest=sha256:d54251618693f2e2d149d2a7318a8d14452ab83726f5756e93efabc21a3123dd

Observation 3c39f7e8-b013-4221-8c2f-f992ec631666 · outbound

This paper cites Lxmert: Learning cross-modality encoder representations from transformers, 2019.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Lxmert: Learning cross-modality encoder representations from transformers, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.599235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.089071Z digest=sha256:297cb86ded9de406cd6758e1e53bb4486368938727f7545216d0b1373a851b9b

Observation e6834180-8224-4b54-951f-a16af1736e98 · outbound

This paper cites Uniter: Universal image-text representation learning, 2020.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Uniter: Universal image-text representation learning, 2020

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.441957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.145387Z digest=sha256:328126311e375385d97d5a93276f0b9a6e7087a1f1cf8eb80ceb78721dab7d2c

Observation edf10535-f2d9-4a8d-af74-9a019feafb9a · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learning transferable visual models from natural language supervision, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:00.211295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:00.211295Z digest=sha256:190e02ad41e79778fdcbab66376a5bf0ec53e41fdd79157b20bf2d944223a78b

Observation e307d979-e9d1-4bfc-a8de-6c5589e7d12e · outbound

This paper cites Le, Yunhsuan Sung, Zhen Li, and Tom Duerig.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.262751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.304151Z digest=sha256:37f592f9b0b215bc5f4a465665a3c0cce665e28dd08456662152a0cae57c8a19

Observation 0165464b-00b6-4e8e-a48f-96a31ff68147 · outbound

This paper cites Evev2: Improved baselines for encoder-free vision-language models, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Evev2: Improved baselines for encoder-free vision-language models, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.072117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.398803Z digest=sha256:c82e1e992be07983fa0d2d635083f0357deec6a7d75a46120a214af3080be9d2

Observation e2ff3864-8d3d-4561-aa19-d46c5f3a310c · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.866112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.520998Z digest=sha256:c80e319a60534c811f147106e3b6a21c5562c92a93c0b7e51fc7d36c4cdc4d8f

Observation 95b64913-819f-4659-ad65-d74794211a08 · outbound

This paper cites Introducing Gemini 2.0: Our New AI Model for the Agentic Era.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing Gemini 2.0: Our New AI Model for the Agentic Era

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.721855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.644759Z digest=sha256:74509de12aa6c5acdaf78b92dfb69ec185977be14385cc0742bd2f27fcc57729

Observation bdf30c32-9d80-4129-8d5f-a9364fd3e3b0 · outbound

This paper cites Gpt-4o system card, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4o system card, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.594072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.733515Z digest=sha256:c5f0e0d20d3057dafeca760252725c099e1df57f0ad41558563c1399225c75a9

Observation ab5f2dc8-7eb7-4aff-80dd-9076505a8771 · outbound

This paper cites A diagram is worth a dozen images, 2016.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs A diagram is worth a dozen images, 2016

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.426138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.823620Z digest=sha256:a465279aaf3369c809838e2cb1a445c3553bca947a44ffb6c7f0a5d9fcf1eb72

Observation d901f0a3-432b-4f3b-aebf-19d4b261882e · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ocr-vqa: Visual question answering by reading text in images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.206904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:00.934581Z digest=sha256:cec3a429d29467a8e199b0703e45520e0c78e1618d63feca4577dd8faf23fcca

Observation 3f998ecc-864c-4531-a960-a465f0fd37ed · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.060454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.029385Z digest=sha256:19e3639d38c232341a7ea747dc0702ce7ed444332ad4731bebf21c1c31256718

Observation 08c6473d-8e3c-4986-b1da-3893616464c5 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.897367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.115297Z digest=sha256:e87ccbcd562017e7bb4fa3ecf37ce14dbda162feef656221eab9c0f276ef09ac

Observation bf849b6d-5d5e-4054-a099-03bda40f6d56 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.212066Z digest=sha256:e1a20f2dd0558b0d719b4e903e8c245b8020c324ad4718334595be7b7e135c82

Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:01.334554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:01.334554Z digest=sha256:22d140bdee20aa9338b93dc68e538db6cc29d3e3a48e69cd040b957d677b7649

Observation 7df63926-5e1e-41ba-b9f6-9a9c46e755e6 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.610441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.400126Z digest=sha256:c784cdb83a9604de60a6af97a45f9d1554f224e6b502d6a5776828fe3c8695ae

Observation 68aa9c35-5de9-4e43-9aee-38970d0c5ab7 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Measuring multimodal mathematical reasoning with math-vision dataset, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.468702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.467755Z digest=sha256:1b3925322b34a6b0551a340288835c47311457d299054467fb2c5620ff8e9878

Observation 6b2e37e9-dc90-4229-b1fd-7aba4120cf33 · outbound

This paper cites Claude-3.7, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Claude-3.7, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.326033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.540750Z digest=sha256:fb0372748bd372d6fde0e18598c4b120e37c8c402c37337f4e466a2d95ae6991

Observation 510321dc-36d3-4e2a-9f78-94bb6aa88056 · outbound

This paper cites Qwen2.5 technical report, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2.5 technical report, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.178864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.641120Z digest=sha256:c11849b1662a7c862883b48e3089f3b3c720d83daf4d27a587108ba12d5ba95c

Observation 170d1a4f-c702-4868-bd5f-03b805dbf84e · outbound

This paper cites Mineru: An open-source solution for precise document content extraction, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mineru: An open-source solution for precise document content extraction, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.014990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.733700Z digest=sha256:abf7ff1220a9e71a618697818786dd6fda22163a0704ea8e590314c53dccb22f

Observation 4c60abfe-3b64-409d-8ce5-e5c6e6f7704e · outbound

This paper cites Deepseek-v3 technical report.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-v3 technical report

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.843258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.833051Z digest=sha256:6197784ed014fb832c8557331712c6161179cb8681bdbb1901666814617f1545

Observation fef5838e-27b0-4f96-a1fc-0e97b6ef2156 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.706057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.909165Z digest=sha256:53ee47cd36f8e162845946533d2edccc66182249d42386b17e6fb3c46bc18358

Observation 85c4e7d7-c828-42a9-822b-530d57871178 · outbound

This paper cites Introducing our multimodal models, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing our multimodal models, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.449469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:01.994475Z digest=sha256:4b2c76a17be91a5a684f6aea2de6c20a8758d6fde2bf7a6c4123459d36d56591

Observation b1515deb-5dbb-41e9-aa5b-4c50b36a65f6 · outbound

This paper cites Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.254171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.078331Z digest=sha256:f5f37515c3d22f29252121968d223bf0917f89b75c897c2436fccb72b4f1dc84

Observation 288caf5c-3887-43dc-a5c0-798a40f0870f · outbound

This paper cites Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:02.241977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:02.241977Z digest=sha256:0293b5cbad89375abdf96eb4e2fa8cc6bc8c0755171a42b7623b2a16529de451

Observation 132f83f7-0961-4137-8bbb-9dbd4ccda514 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:05.072111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.315018Z digest=sha256:56244c6e462b0fcbffbb10a0aae5f4dd321fdfa29740450db2602064a290e91d

Observation e3752dbd-0ef9-45a2-9bbd-02c598537319 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.921260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.379786Z digest=sha256:314e0564b30c027f5de8d0b664cb9fca5a4b3856a8df745cbf9d41e45ec13dee

Observation 0533e597-c71e-438c-af5f-70a61bf93df6 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Building and better understanding vision-language models: insights and future directions, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.758949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.482049Z digest=sha256:b5c415fa6baf538c544e10d53ee9ee57d858e29fb3d47799ed055d38de3e51e8

Observation dba9b979-683c-4a3f-baa4-df085e00bf06 · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Improved baselines with visual instruction tuning, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.535501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.594606Z digest=sha256:436815f6b7779b70aa4e1ce146c7203ae18c37c4811b2293a0ad8cadb438e633

Observation 34071225-9ce5-4881-8e4e-cdaae9a96577 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:04.279722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.668534Z digest=sha256:9b77d7c804df3aefb5988da66140865d73d56fe9c3712908a866d6b29408a784

Observation f8920688-1ab2-4d0b-a3c3-4934be1addb0 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qvq: To see the world with wisdom, December 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.124055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.742793Z digest=sha256:caebce0acbf9db9bbd064d786d25bb6389e765737f9c3375df29cb1cd375c292

Observation eb4b8d66-110c-4d78-81af-ba890d089ee8 · outbound

This paper cites So the final answer is \boxed.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs So the final answer is \boxed

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:03.865748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.858561Z digest=sha256:5eac241ace8dea12db16604d4ffbb915e5d49b0e36e8258f0670b2fe886b54e1

Observation 1724ab91-8c2a-4e98-b23e-3b69f20ed704 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:03.646708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:02.974741Z digest=sha256:95d93db7baea160431641fb05317cc157aa29a0509578d8fbb1e0e778b4f22ca

Observation 8fd73f73-ea3a-4c2e-8e8b-760cf06a0ba5 · outbound

This paper cites No," please identify the main unreasonable aspects or obvious flaws in the solution; if.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs No," please identify the main unreasonable aspects or obvious flaws in the solution; if

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:41:03.400314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:41:03.076855Z digest=sha256:e146a6b9427dff1e8c0938d163d743b2561f45c305b94ee4a838da2400019655

Pith citing papers

Observation 135949bd-3881-4ccf-9c25-8f809311f889 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.671614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.671614Z digest=sha256:8ebcca57942f325c528c104cd747de80db2c89a1a8693da7a839eb6e2deb8197

Observation 429530aa-4830-4ffc-a5c8-fdadb6c92953 · inbound

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning cites this paper.

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.166191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T19:21:30.235583Z digest=sha256:fb1d526a568591a720e9fd1f1575fae56150bc35fa9c4a1a8b295083c0e5550a