Pith. sign in

Paper Citation Record · LEDGER

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2406.05967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05967 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:22.651869Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 79047972-252a-4eca-b38c-ed69d7866932 · inbound

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages cites this paper.

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:07.673361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:07.673361Z digest=sha256:bc4b728511981d0576b0154b0184bde2e9ea07c823fe5f5d357029ebb9b9c1d7

Observation 8c372f18-98a6-480a-841b-83496213dec2 · inbound

Maya: An Instruction Finetuned Multilingual Multimodal Model cites this paper.

Maya: An Instruction Finetuned Multilingual Multimodal Model CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T19:10:57.451130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:10:57.451130Z digest=sha256:a4d21abfe4eabbf6d207ad74a85f0caad120db10756e636c2c870615a3daea28

Observation a91608e1-9925-42b6-a46b-87b2041d3e67 · inbound

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries cites this paper.

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:36:03.809020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:36:03.809020Z digest=sha256:314c146641fafe58de9dc7c7bca0f283bc8b955fd10298e94f2d96f1a1b6f67c

Observation 65c202e4-0623-4608-a79f-c9f50c3d4afd · inbound

Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model cites this paper.

Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 2591

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:55.310900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:24:55.310900Z digest=sha256:bb716275780f283243363ffc7e03311310af6194fad27bff96b3a442cb1e8c37

Observation 5820ac62-96ca-4345-8348-a0c35a8c8280 · inbound

The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes cites this paper.

The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T17:16:19.303266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:16:19.303266Z digest=sha256:4c79ea3e818969cbea9b7928bdd879a6ae65d6b6c1f01f71112083ce5753485d

Observation 31db25bd-bb54-4708-a346-29a4de1a48de · inbound

RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation cites this paper.

RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:44:39.564015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:44:39.564015Z digest=sha256:eb19ba20ea7825c26bc267ad936a97fc0009b0f4a9d4d2ad0c87c29ce4aea83e

Observation 47ede23a-17a3-4574-bb5f-669b1d7e3c49 · inbound

Behind Maya: Building a Multilingual Vision Language Model cites this paper.

Behind Maya: Building a Multilingual Vision Language Model CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:22.651869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:22.651869Z digest=sha256:342471bd7ea2c64f849913ac2323170bd432b3b0fe4d61a518787ec150bbfb52

Observation 9da75c24-78c1-4c9f-a009-e9e5cef09de6 · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:46.670649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:46.670649Z digest=sha256:d11464649064acf7135a841c6f13941d2430a3dd649df1d217e5b4e30474f976

Observation 8b9cd554-a0f1-4e01-b5c4-0f9776fba7a0 · inbound

CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention cites this paper.

CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:38.873716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:22:38.873716Z digest=sha256:3dadb684d0a307e2760975d00355ff8f5e467ac78efcb0e91089dd22c7e289c4

Observation 83d00571-df45-4872-9d45-ebc8ba62f3cd · inbound

CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation cites this paper.

CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:23.308048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:42:23.308048Z digest=sha256:f61df89cec2311395404d36d7a3b119bd7610a954c6a3bdf73cac581e5f7daed

Observation 175b67bb-690d-458b-ad3e-bc16960eb334 · inbound

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models cites this paper.

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:04:23.217835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:04:23.217835Z digest=sha256:cec2335aaaba0403ce0f5b1c22749d882d1c31b00ab5979ca40cca199e0ccffe

Observation 29b33775-fc13-45a1-b809-616da9dad5de · inbound

The Devil is in the EOS: Sequence Training for Detailed Image Captioning cites this paper.

The Devil is in the EOS: Sequence Training for Detailed Image Captioning CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:07.433172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:07.433172Z digest=sha256:dd5660cc165ff61bfba61ec623a0a8120882f907d8344d8dfe13a22349a4610f

Observation 9618a300-00c2-447a-aeaf-eb6c98d8ef6f · inbound

Grounding Multilingual Multimodal LLMs With Cultural Knowledge cites this paper.

Grounding Multilingual Multimodal LLMs With Cultural Knowledge CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:51.204793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:51.204793Z digest=sha256:dfd249c22b30a5014932457f67003c31a4fdbab6f479ff9264abbb4eaa890857

Observation 428e2e7d-8199-4a49-9c58-b2f70941b2a0 · inbound

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding cites this paper.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.200780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.200780Z digest=sha256:80b193077639888b22392282800156a28ae98283ca32188f86099743f010f568

Observation 91ecdc94-518b-434d-a4c6-ac1ebc6593d1 · inbound

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly cites this paper.

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:06:50.173571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:03:26.144452Z digest=sha256:10af818c937bf059b41b6f072e06ecfca8993abc215674438efa76efe7d70462

Observation c044eadd-ccc8-4f7d-bbd6-d9eca4e12ad6 · inbound

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages cites this paper.

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T11:21:57.622339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:21:57.622339Z digest=sha256:8b103fed96d99f46ba9848ad40cd2068b768646317134650486eea614e131f5a

Observation b8b02e51-595f-48e4-9ff5-bbfe480b8c77 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:19.995920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:ed2bf97c0d276241e09bb751c9242c608be5f3e4329cf349130d379f0a0ddba6

Observation ed2ff688-0306-4f92-8ec3-2d0b41d5bb2a · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.742039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:03ebcb3731f8c1f6fddb985cef30069571bd5d0560511e77b087ad5b47d3754d

Observation 463be97e-c7f1-4ac8-b42d-51fac10554f7 · inbound

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs cites this paper.

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:08.858876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T10:34:33.399684Z digest=sha256:df222df39daac287585ab8dbca703efa201a47028ead7df538ba6ee2d123b0b6

Observation 69095852-e8a9-4c98-859a-4f6d520332ff · inbound

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs cites this paper.

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.698921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:49:23.895938Z digest=sha256:635520b1aafce2cba83ed31d517cadee5edc91c6bd1fd26411468a60851c7035

Observation 10c268a4-983b-4eed-9c71-46b3c1a7ae37 · inbound

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models cites this paper.

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.252812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-03T14:48:06.514693Z digest=sha256:2773431a8e4d47191845263ad214416e329d5a19eaffde6b04cc89a1de8b99d4

Observation 34d2eab2-c807-403d-a919-86b5c21fa622 · inbound

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation cites this paper.

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:28:29.511963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:28:29.511963Z digest=sha256:ff4d040930e2fc0f4f2b9a219eae3948e7e408742d7d423934d7979609e24982

Observation f7681ec2-f2ad-47ef-8814-7fa80d1444a0 · inbound

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL cites this paper.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.540281Z digest=sha256:b5f7fb2462792ee78ae59f297a0239c7096b6f268481b9a19938ae3e2d6af1c0