Pith. sign in

Paper Citation Record · LEDGER

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

As of 10 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2507.13405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13405 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:41:13.655335Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:01:48.251493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.200070Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ecf877d-c8c4-4b1a-8c1f-99a4d8bd3e9c · outbound

This paper cites GPT-4 Technical Report.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.166689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.166689Z digest=sha256:132935d11345275d569baa6e198391d066c61bbffd6956794d6866f79c4092e2

Observation 39f49f9f-7fe3-4f23-afa7-7c5ebae49050 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.409602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.409602Z digest=sha256:e00538a687a32aef01c2052c93b4b5867227e4777c8bddc134f1c4186b875e72

Observation 284bbb27-812a-4401-b4af-d444e8cf7cde · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.639417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.639417Z digest=sha256:54c1881661446d006ef9b560a74d713c6f935ee35c69175c368211bceeed7fc4

Observation bad059c7-b2ee-4e1c-ab95-768d91d5a10d · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.716870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.716870Z digest=sha256:67938c13eec42a2afd240df4b58659c7ed7d5e37c2f092a171fa5888ce1fa19f

Observation 1720286c-ee3f-4044-bbb9-3943b9585fb8 · outbound

This paper cites an unresolved cited work.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:41:14.859133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:41:12.809200Z digest=sha256:52db055b182d1d235e137c7735eb8bf1882f63841a985b6241a401d5c54b620c

Observation 66eaf9cf-7123-4b91-bd8d-f881b07038ea · outbound

This paper cites NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:41:14.160149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:41:12.884118Z digest=sha256:84b0968a9d0b62ab4f10008832c0fb9c58b08864cd4e49fe2f5605e19693cf6c

Observation aeac34aa-ad93-4f30-a0da-a40b4b5066b8 · outbound

This paper cites VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.007547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.007547Z digest=sha256:4a681cd91e16a0d71cb11bdcb696412a523de60fe38325e82e722716bcd3df59

Observation 13202050-d695-432d-b7a4-79bdccf119ae · outbound

This paper cites CrowdHuman: A Benchmark for Detecting Human in a Crowd.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark CrowdHuman: A Benchmark for Detecting Human in a Crowd

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.138030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.138030Z digest=sha256:ab057782b93996fc86342353c52c3f7f493c2128dac91c61bf121349a6303837

Observation f3215973-2793-40bc-82db-132191999531 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.278562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.278562Z digest=sha256:9ba72b09c0ac2c5ad2378afcf3090b068a6fdc8cb3dccf006d2029feb3b42212

Observation 49b1b18b-a872-47ff-8655-2e2f8a63c8c1 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.357316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.357316Z digest=sha256:3b69eeebb12342ebc355515df54dd8c7ffe2817af4fb1d7c20db1504c1098c17

Observation 3cb4cf62-8367-4010-a1de-5c6103ea2db4 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.436572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.436572Z digest=sha256:c793fea84cb7654350aa463232c7079b934df92693eb62036bbfbd171629d2b7

Observation f0c63bf0-9da3-404b-a112-d41aa3f3a32e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.513227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.513227Z digest=sha256:296bbfa2caab6ae3118d36867ed1ede8123d296f147627991cbb4452e1325627

Observation 09201af8-cbf8-45e1-bfca-4b3e0b1d79d0 · outbound

This paper cites When in doubt, use more conserva- tive qualifiers.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark When in doubt, use more conserva- tive qualifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:41:14.620596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:41:13.590449Z digest=sha256:2b631651e05dbd96c32149ff71480630d334a64fce2f813fab5b164f0b6fb57e

Observation 28ff9e58-16a6-4a6a-beb9-2db731eed007 · outbound

This paper cites COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:41:14.459868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:41:13.655335Z digest=sha256:8d6573d784f26591b3c36ca62d99021e2325d240524911e02cfe50be5d0f768f

Observation 94ee7c5c-6f97-4031-9c8b-b618516ea93f · outbound

This paper cites Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.234673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.234673Z digest=sha256:26bf021420dc11b00e6e1b0581887e8124df9a7712fd066a023555cfac06f9de

Observation 01a7d63e-f7dd-4c42-8484-267f692b0260 · outbound

This paper cites M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.203791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.203791Z digest=sha256:e0fb132749fc1a2573c794e0411fa64f56c288c3af78f14af7dcf12fa3257091

Observation 8c717dd3-edca-4384-b167-0a529f19c289 · outbound

This paper cites R., Bashir, S.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark R., Bashir, S

Reference 2021

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:41:13.982818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:41:13.076746Z digest=sha256:cfca878f3790b00e0821a2fef22999c91d497f7d11383737150111c349dc73a8

Observation 3d466147-a3f3-4582-81cc-cf2bfb5d755d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.553911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.553911Z digest=sha256:606285d45ea20debd62f2d1bb90ab59ab9d841d78bb958f8fdd7dc5317cf9d8c

Observation 97e4749f-3741-47c9-840e-643562c7864e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.321982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.321982Z digest=sha256:6b3a40eb3f7be06b6b77f8ab39cb49dcae8b948c3a86f8f9f56874241a58e2db

Observation b305033a-7bc9-47a0-86da-c716cee83a57 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.474596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.474596Z digest=sha256:68ce1d81e7d0c663046f9928fa3079f14b04b97ab6a4fd3b835e5c8a972a4a39

Pith citing papers

Observation 4340bbb5-e9ee-4e65-a263-2a16a363cfbe · inbound

Scaling Mobile Chaos Testing with AI-Driven Test Execution cites this paper.

Scaling Mobile Chaos Testing with AI-Driven Test Execution COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:48.251493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:48.251493Z digest=sha256:163830bbfd3a4296494a5f2adf8c073273a1282e2730ab0dd0311e6276aff0b5

Observation b651095a-580d-412d-a407-a5cb8419b22e · inbound

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling cites this paper.

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.201757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:32:41.348880Z digest=sha256:536e90a7784562b4e95e7ab1b8fce36a69eeb96d998452b9a6b55ffe4c3fb15e