Pith. sign in

Paper Citation Record · LEDGER

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2507.13405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13405 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:41:13.655335Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:01:48.251493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.200070Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ecf877d-c8c4-4b1a-8c1f-99a4d8bd3e9c · outbound

This paper cites GPT-4 Technical Report.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.166689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.166689Z digest=sha256:44a23c6e6ff0b4dd6d7a4bc9790b13bd2b0e7eb0285351173f7b73fbf344ed00

Observation 39f49f9f-7fe3-4f23-afa7-7c5ebae49050 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.409602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.409602Z digest=sha256:677aac09fe86964211df55ec9890bda3d834368eca3aa8a51d2e27e1e28d97d6

Observation 284bbb27-812a-4401-b4af-d444e8cf7cde · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.639417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.639417Z digest=sha256:fdbfc3589e65e9cbbb9279823fab49a3d638330ffbf55bb8c53eab174a8895b7

Observation bad059c7-b2ee-4e1c-ab95-768d91d5a10d · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.716870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.716870Z digest=sha256:072f66b866698b54764c216b100d3501f876932675b195d3a6f64d964bacb06a

Observation 1720286c-ee3f-4044-bbb9-3943b9585fb8 · outbound

This paper cites an unresolved cited work.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:41:14.859133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:41:12.809200Z digest=sha256:55b21553eefd15004f699b6447bc2d95deaa3cac4d9b142f1a56cb3875fa504f

Observation 66eaf9cf-7123-4b91-bd8d-f881b07038ea · outbound

This paper cites NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:41:14.160149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:41:12.884118Z digest=sha256:87500cf10713fb11d860b4c8b21638ef792156e8b62afc08d324c2e951b8414f

Observation aeac34aa-ad93-4f30-a0da-a40b4b5066b8 · outbound

This paper cites VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.007547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.007547Z digest=sha256:533e6bf69ad17f92dd7fa5549478c5398429e482193e17f44114c1c24df7ad1f

Observation 13202050-d695-432d-b7a4-79bdccf119ae · outbound

This paper cites CrowdHuman: A Benchmark for Detecting Human in a Crowd.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark CrowdHuman: A Benchmark for Detecting Human in a Crowd

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.138030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.138030Z digest=sha256:03002f220ac5c2b04cfbb7a3e6fbad8f900c5442995ce193f65550d73e64bb9c

Observation f3215973-2793-40bc-82db-132191999531 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.278562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.278562Z digest=sha256:7a930de93d38d4830d698a429dda7a55559c1e1e9a52c2132e85c1587ee3304c

Observation 49b1b18b-a872-47ff-8655-2e2f8a63c8c1 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.357316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.357316Z digest=sha256:b2e667dfb2c303726376685e896dc1120dad340e5d4d08bc24372c7b5cab6476

Observation 3cb4cf62-8367-4010-a1de-5c6103ea2db4 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.436572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.436572Z digest=sha256:871de6ff517b98f12911246cad6dfa94e9eecda5a38a32960218eb08c8c1196b

Observation f0c63bf0-9da3-404b-a112-d41aa3f3a32e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.513227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.513227Z digest=sha256:0a5bac163c83507d6f510662c7b1ee351fdf540353e4d465f1f10037476b5b21

Observation 09201af8-cbf8-45e1-bfca-4b3e0b1d79d0 · outbound

This paper cites When in doubt, use more conserva- tive qualifiers.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark When in doubt, use more conserva- tive qualifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:41:14.620596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:41:13.590449Z digest=sha256:c64b4bbf8eda03c21e4a85765d08cb67713fe1f2371148bf38af8f0ecbe4d800

Observation 28ff9e58-16a6-4a6a-beb9-2db731eed007 · outbound

This paper cites COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:41:14.459868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:41:13.655335Z digest=sha256:7891272227b358aa6c1f3d1b5c0868414878c0767caec45ffc21aa766e01817a

Observation 94ee7c5c-6f97-4031-9c8b-b618516ea93f · outbound

This paper cites Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.234673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.234673Z digest=sha256:1fbc9f728d3071fdaf21bdff908bca7fff4b41fd39810fbe218a895f68b90607

Observation 01a7d63e-f7dd-4c42-8484-267f692b0260 · outbound

This paper cites M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.203791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.203791Z digest=sha256:986036ea7dba84d6d0dd856b9c74c80ca638678618529e149239a0b832f544da

Observation 8c717dd3-edca-4384-b167-0a529f19c289 · outbound

This paper cites R., Bashir, S.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark R., Bashir, S

Reference 2021

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:41:13.982818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:41:13.076746Z digest=sha256:80912781cd1daefa2dc3176e7d51b3ebe527f113e60a5f2e91249ff3ee1d2a6d

Observation 3d466147-a3f3-4582-81cc-cf2bfb5d755d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.553911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.553911Z digest=sha256:e86a6402802a43db7ad75cd82e0c2e71a293410add2c4973576ef825702e6661

Observation 97e4749f-3741-47c9-840e-643562c7864e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.321982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.321982Z digest=sha256:8f021f971837362510595dbec118e11401ff6e32b1f27fa40045e5b43fdec7bb

Observation b305033a-7bc9-47a0-86da-c716cee83a57 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.474596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.474596Z digest=sha256:bb70594a27c5df905ce6a6ce4fae86d34abeffe330e51ec44cd87ed19d0bf2ae

Pith citing papers

Observation 4340bbb5-e9ee-4e65-a263-2a16a363cfbe · inbound

Scaling Mobile Chaos Testing with AI-Driven Test Execution cites this paper.

Scaling Mobile Chaos Testing with AI-Driven Test Execution COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:48.251493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:48.251493Z digest=sha256:0e96012da5d2ac679d0574d26f314933e6c3837de5bfcddfd87074f5af4e00ff

Observation b651095a-580d-412d-a407-a5cb8419b22e · inbound

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling cites this paper.

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.201757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T18:32:41.348880Z digest=sha256:1721b626e71b7b54dfa9cfb438dc0d9f3925e38d802ed04820500ab4827d24cc