Pith. sign in

Paper Citation Record · LEDGER

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2505.24120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24120 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:03.076855Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:03.671614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T05:55:31.163509Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7584910-2c37-4e9a-b185-661879891cde · outbound

This paper cites Gpt-4 technical report, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4 technical report, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.070452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.070452Z digest=sha256:75841334fcea01ba30aef10d907096326685e7d1dc70cce3a36561c20c857288

Observation c7f311c7-8fad-4b1b-8402-395389ae472a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.813861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.118262Z digest=sha256:01750e858761267f335620a01197b7047b70b7957331c345d3fb076a7acc7ee3

Observation 89ecddcb-d994-480e-8a5b-2eabb2e6abc4 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.703189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.202715Z digest=sha256:82c546df0cd21ecc01c59b449b694e4e60b4fa99aa8143c1f5d612c030390357

Observation 94fc0825-f2b2-4750-b2e2-1a5f9fdb828f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.273403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.273403Z digest=sha256:dd6a1646f3ca1fcf9690bd34a5b19dc1d621e4d2b18656626705c95fa54de746

Observation 30fdd865-079d-4720-a636-45aff8245f09 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.348853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.348853Z digest=sha256:ee216d6ecb899e80969bec6584b9e5dff367ead0858dc9ed6a95396ba1838fb2

Observation c7d9aefb-fa90-4664-b2b6-5d030929f230 · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Kimi k1.5: Scaling reinforcement learning with llms, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.582667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.418298Z digest=sha256:25ec803e9452446a43532256f2c623fe1ba965915362d4e30bd5f5835bec7fd5

Observation 161e185d-2897-4d8b-a607-9c89c2732e88 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gemini: A family of highly capable multimodal models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.441409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.496657Z digest=sha256:ac5224378852edcaa7f9b8288ff6373962d2acfc975bda3b13f5b181a69c38d3

Observation 324b1f69-6911-4cc9-b3fa-04e331ba0a3d · outbound

This paper cites Hello gpt-4o, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Hello gpt-4o, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.301657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.588916Z digest=sha256:03e65850196b6a7fc78e5f28481694998a3ecc47f64f66f2d368c0556efff11f

Observation 4b54aee7-b2fb-4eac-b9a5-a85a64adee3d · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.154599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.657293Z digest=sha256:73cd22e5ccb10e7e74989ef0a2386d0fd4d4a3fa1167c7ed77315a49b11fddd5

Observation 3b68a956-1007-41cf-af05-323095274103 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:09.028912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.737185Z digest=sha256:fad3249a5edb514c0cb3c80ce669b35ab8aa48048a7095d7d43714156c8d6625

Observation 691bd1e8-04a8-46f7-b709-c11d53c7ca57 · outbound

This paper cites V Jawahar.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs V Jawahar

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.888367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:40:59.885061Z digest=sha256:b3a5485f42321239a33e736c0dd9e720a0c144e8a7201021304767062ff46497

Observation 5ab11a5d-6aa8-4048-8acc-13bb2133af01 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.731263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.011399Z digest=sha256:a6fd0c5fc7102c076749fc2cbb88fc8576a8db786ddc15507284df7d9580dede

Observation 3c39f7e8-b013-4221-8c2f-f992ec631666 · outbound

This paper cites Lxmert: Learning cross-modality encoder representations from transformers, 2019.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Lxmert: Learning cross-modality encoder representations from transformers, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.599235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.089071Z digest=sha256:56995f408e084d24e83b90d453815a85bc1246dabf57c60afa7fd6f6aa9f1843

Observation e6834180-8224-4b54-951f-a16af1736e98 · outbound

This paper cites Uniter: Universal image-text representation learning, 2020.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Uniter: Universal image-text representation learning, 2020

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.441957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.145387Z digest=sha256:80f45021f0873eeb45a899fa98771b56b3a5ed5ca15cc6ce2cfa60d18186f088

Observation edf10535-f2d9-4a8d-af74-9a019feafb9a · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learning transferable visual models from natural language supervision, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:00.211295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:00.211295Z digest=sha256:190e02ad41e79778fdcbab66376a5bf0ec53e41fdd79157b20bf2d944223a78b

Observation e307d979-e9d1-4bfc-a8de-6c5589e7d12e · outbound

This paper cites Le, Yunhsuan Sung, Zhen Li, and Tom Duerig.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.262751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.304151Z digest=sha256:8055da8a0168b64c6d34a85967acca20171c460ae472144e4ae0acafe030c95a

Observation 0165464b-00b6-4e8e-a48f-96a31ff68147 · outbound

This paper cites Evev2: Improved baselines for encoder-free vision-language models, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Evev2: Improved baselines for encoder-free vision-language models, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.072117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.398803Z digest=sha256:592542a1657ad69d9a4c1ffe5b33b2b02914d75c629362f636ff45136b49cad1

Observation e2ff3864-8d3d-4561-aa19-d46c5f3a310c · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.866112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.520998Z digest=sha256:37524662dc09b12dd78c9bbbd7840765d8a6ef74fc6f5cc3ac8a62a9c45ebf5b

Observation 95b64913-819f-4659-ad65-d74794211a08 · outbound

This paper cites Introducing Gemini 2.0: Our New AI Model for the Agentic Era.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing Gemini 2.0: Our New AI Model for the Agentic Era

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.721855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.644759Z digest=sha256:7d5c2e08ab823a908ea728ddbdc7e4f50fe45e6e3ef4f14c28df7715e5b4b532

Observation bdf30c32-9d80-4129-8d5f-a9364fd3e3b0 · outbound

This paper cites Gpt-4o system card, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4o system card, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.594072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.733515Z digest=sha256:8f33921d427085a433ef2a3eb0ab2cc0da9828767b9f358d6d7684eec29ed59f

Observation ab5f2dc8-7eb7-4aff-80dd-9076505a8771 · outbound

This paper cites A diagram is worth a dozen images, 2016.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs A diagram is worth a dozen images, 2016

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.426138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.823620Z digest=sha256:f1b6eb47503ec6cef75c0f39db8f1df4b4563a7cba6748d06d8465c100b908ae

Observation d901f0a3-432b-4f3b-aebf-19d4b261882e · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ocr-vqa: Visual question answering by reading text in images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.206904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:00.934581Z digest=sha256:e099e3a5eea63011be37072b004f39bbee5326dbd527272c6057c83e1c7b7b2a

Observation 3f998ecc-864c-4531-a960-a465f0fd37ed · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.060454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.029385Z digest=sha256:61c4caac0cc5b01db32e97965d1dec7c83d291e5958e54c715900fbad915911f

Observation 08c6473d-8e3c-4986-b1da-3893616464c5 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.897367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.115297Z digest=sha256:be6459674f0d73ca1a8405b6f1642d9208a97a8b79b05718c61d82a97010cec2

Observation bf849b6d-5d5e-4054-a099-03bda40f6d56 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.212066Z digest=sha256:7c4e336a6f1cd1a710f9daa35f949e1d04a3915faf4812ff24f5ec3f3dd498bc

Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:01.334554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:01.334554Z digest=sha256:c84ef2b37942ab8b92e2ec863bf43fa26e1755248938afa01eefb02bb2b9204e

Observation 7df63926-5e1e-41ba-b9f6-9a9c46e755e6 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.610441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.400126Z digest=sha256:74eb25cb988c983f91bd863e3b5747b807ab3b28190a22617fc0f5f0a77273d5

Observation 68aa9c35-5de9-4e43-9aee-38970d0c5ab7 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Measuring multimodal mathematical reasoning with math-vision dataset, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.468702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.467755Z digest=sha256:77cef50f4e79244cee84c2644728ddaf2af57373432f8c0b6a49c9bf972d4802

Observation 6b2e37e9-dc90-4229-b1fd-7aba4120cf33 · outbound

This paper cites Claude-3.7, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Claude-3.7, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.326033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.540750Z digest=sha256:08f73927046785a26fb6044dfa84e5068af6ac08732f8a9f591daa008fc7c33c

Observation 510321dc-36d3-4e2a-9f78-94bb6aa88056 · outbound

This paper cites Qwen2.5 technical report, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2.5 technical report, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.178864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.641120Z digest=sha256:af9548950209fa758ee7dceaffad26931bae1cbd88f33a298b0f7a5f5006333a

Observation 170d1a4f-c702-4868-bd5f-03b805dbf84e · outbound

This paper cites Mineru: An open-source solution for precise document content extraction, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mineru: An open-source solution for precise document content extraction, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.014990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.733700Z digest=sha256:eaea733e3be599880cb681d6ff4789ebeb7a3160a70ebc1f744fb4ef4bd476b3

Observation 4c60abfe-3b64-409d-8ce5-e5c6e6f7704e · outbound

This paper cites Deepseek-v3 technical report.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-v3 technical report

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.843258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.833051Z digest=sha256:9e32e2c5717ab87a851c82e6330311ee11cbd51adcec4cf9ddd093dcd15fc3ff

Observation fef5838e-27b0-4f96-a1fc-0e97b6ef2156 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.706057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.909165Z digest=sha256:0b96488da9b38660dc4994c55eac6a4f4c80c8e3041cdfc775ec6fbed1253078

Observation 85c4e7d7-c828-42a9-822b-530d57871178 · outbound

This paper cites Introducing our multimodal models, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing our multimodal models, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.449469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:01.994475Z digest=sha256:009cc7224f37c59172868df2d158e7f647062aedc4e9d5d8c21acf8d4c9773f8

Observation b1515deb-5dbb-41e9-aa5b-4c50b36a65f6 · outbound

This paper cites Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.254171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.078331Z digest=sha256:ce096416d5b9b28337f6ee1b75b278475e99b9b1eece2183ecf481efa75b4b38

Observation 288caf5c-3887-43dc-a5c0-798a40f0870f · outbound

This paper cites Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:02.241977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:02.241977Z digest=sha256:0293b5cbad89375abdf96eb4e2fa8cc6bc8c0755171a42b7623b2a16529de451

Observation 132f83f7-0961-4137-8bbb-9dbd4ccda514 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:05.072111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.315018Z digest=sha256:46be11de99b450a935486a4ed7d21324e0cbb578c9e41cc92ddcb71ab5eecb55

Observation e3752dbd-0ef9-45a2-9bbd-02c598537319 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.921260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.379786Z digest=sha256:2656aea8b865570a25f46a078d2a0213958f16dc9210e866897073c1c3b21fb4

Observation 0533e597-c71e-438c-af5f-70a61bf93df6 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Building and better understanding vision-language models: insights and future directions, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.758949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.482049Z digest=sha256:868b3fbff9c9ef7a3f937467cf17ed4f62f5124f45a39b17a23e1841816959e6

Observation dba9b979-683c-4a3f-baa4-df085e00bf06 · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Improved baselines with visual instruction tuning, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.535501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.594606Z digest=sha256:0957c3fe1a94f3304849b03828a78357ed64df660717c81dfd318a99fb880e3e

Observation 34071225-9ce5-4881-8e4e-cdaae9a96577 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:04.279722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.668534Z digest=sha256:23301c11b9c49ac175082c8817efa453322782c618a630d92b40631f2814b0cc

Observation f8920688-1ab2-4d0b-a3c3-4934be1addb0 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qvq: To see the world with wisdom, December 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.124055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.742793Z digest=sha256:dbfd6904a91113af75182276eb4fe33e09dbd8d03a36da2fd687bfc109876aaf

Observation eb4b8d66-110c-4d78-81af-ba890d089ee8 · outbound

This paper cites So the final answer is \boxed.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs So the final answer is \boxed

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:03.865748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.858561Z digest=sha256:0905ae0950d038ce4cf376ae30cf782e96cc81889f4eac810c72c6149a5fb204

Observation 1724ab91-8c2a-4e98-b23e-3b69f20ed704 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:03.646708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:02.974741Z digest=sha256:2ec3c3cc586d642306f9a9beb7203ed79e536cfe450658fa897f8f6f2edb14fb

Observation 8fd73f73-ea3a-4c2e-8e8b-760cf06a0ba5 · outbound

This paper cites No," please identify the main unreasonable aspects or obvious flaws in the solution; if.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs No," please identify the main unreasonable aspects or obvious flaws in the solution; if

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:41:03.400314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:41:03.076855Z digest=sha256:ee67981ed8d0af0fad06cd286ac5a0a63df567b2f405d911ffc9730155c5c221

Pith citing papers

Observation 135949bd-3881-4ccf-9c25-8f809311f889 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.671614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.671614Z digest=sha256:1b01162bca6b6bb97a8f5bbd9eade046f22732f8034b75c63d5cf29e1d17b04b

Observation 429530aa-4830-4ffc-a5c8-fdadb6c92953 · inbound

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning cites this paper.

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.166191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T19:21:30.235583Z digest=sha256:2d2d71fb8903e854ecdc167d8a284f17c9ee2343d76f00aa15c255820231e729