Pith. sign in

Paper Citation Record · LEDGER

Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2402.07270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07270 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:45.265067Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T21:51:18.139638Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c6f8ff8-ad4f-4412-a6ae-31560ec03131 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.265067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.265067Z digest=sha256:ad801d3f918b143c95409cfc7095f2140cc026d1f536a2a920ad34e8e23f8efb

Observation b7b9e79c-f634-44c5-b915-a523806e4b02 · inbound

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models cites this paper.

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:32:13.899059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:32:13.899059Z digest=sha256:dbdb073c00760566533165d8e3faa77eb6c6269f9320537d29fd5dd32f8e7b95

Observation 012ea568-86d8-4365-b49d-c5caf9d9490c · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:00.442230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:59:48.877783Z digest=sha256:e6fe6118470eccbc386b341c0fce34820c8043b000030aad154fc6be9196fe93

Observation b9402f12-276b-4231-ad03-5b3f92fe9a13 · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:32:20.285018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:32:20.285018Z digest=sha256:7d75974f2b1cc47bb8e97576bc92735da368216828a96983d6aaf26ebe436cf7

Observation b5974be9-6ecb-44d9-b0c9-9c2e79aa3aee · inbound

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models cites this paper.

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:18.144254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:08:41.452018Z digest=sha256:1f26b350e94e4b85569d5a636d9372a409e59943aa0ebf6286891b3066715ced

Observation b0f029c4-825f-48b4-a21b-d377385af2f8 · inbound

What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks cites this paper.

What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T13:17:55.700527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:17:55.700527Z digest=sha256:27522fbb54740d618853af6a7ea3a190f5a6f80d32c7523867cea93e8868b981