Pith. sign in

Paper Citation Record · LEDGER

MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2405.11985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.11985 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:14:20.479979Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c0cffd4-9796-468c-9105-a91c41822bbc · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.703331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:ae6d7f0029acfbf8b7cbaf20a2548eecfe55db6a653e7c491bf2bfbc41d28d14

Observation 0bc929f0-7cd8-4d2c-a166-8b4f02b75fa7 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.036152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:595211bf27969741dd12f28e204304816fe4abadc59becdd6e0ad7fefca6c0bd

Observation d65ed6ca-8fed-4637-9dbd-badccf0998b8 · inbound

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios cites this paper.

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:14:20.479979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:14:20.479979Z digest=sha256:d1cbd88955e3aea28f2668222a97248f9883ac4d80d62b4714d18fbed226ce36

Observation ebb1b6e8-f7fb-468b-9459-b4bfb34acc37 · inbound

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements cites this paper.

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T15:56:19.626881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:56:19.626881Z digest=sha256:27716c21d79cd7f7cafe8004cc11da0300e58a064d5789b49864f4b7aeb74ccc

Observation 38269d09-3bca-4cbd-8bb8-73503b83363f · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.122658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:af255ae4ba4ebdc2253efcaac73813d5580269ed3b6620e3ad4de9a64b5efa09

Observation 61cd1887-8c50-4ebe-85e3-3ea3c9b7988c · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.101037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:0cc38c6ad784133407391940e7076e7aba93b412ca6761f1b6f4b7f08570704e

Observation b33f1460-4f7d-43ab-8b0d-7e224e99433d · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.096968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:87d90233bf399eeab2ab160ce968d4e1481d2a22f187027f2f7984cf623f4892

Observation 5c42ea2b-b3de-4070-8fb4-ef412deb4b9d · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.648631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:27.648631Z digest=sha256:6fa85f2c089fe974022ce23bc30a56904d915bae1d9693d494386bf86eece89c

Observation 8f0e8020-f461-475d-83dd-13ddd401c8bc · inbound

Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts cites this paper.

Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:09.826103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:09.826103Z digest=sha256:67b01197a3e3d413a9f7f49505bc20756bf32c4ef6e5e9e2945534ae096df46a

Observation 03f0d185-6e7d-4775-a7bf-4f27119e219c · inbound

Multilingual Multimodal Software Developer for Code Generation cites this paper.

Multilingual Multimodal Software Developer for Code Generation MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.185565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.185565Z digest=sha256:025708db3a10ccd58916580f413466ff15642de887200a33101219fcf7a13749

Observation 4ddae192-bd41-49eb-aed9-21f4a513d996 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.775085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:017c835f743a1592df26c250edff95bee45c5b010ddac806561113bc0483267d

Observation fa77b7a4-1134-4c0f-9ab0-63b9d36b2162 · inbound

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA cites this paper.

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.310528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T17:53:08.136677Z digest=sha256:36e5b99e986ec5cb57e4f63a169141ed165f7f2dd8ae6f77abedff77fb327634

Observation 34fb4af9-5580-4886-ad1e-d1a63464aaad · inbound

Multilingual Vision-Language Models, A Survey cites this paper.

Multilingual Vision-Language Models, A Survey MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:02:37.275213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T13:02:08.000814Z digest=sha256:22ac3c93a6c17769c3479eaae18fb4fc217bf7ef86c97f65a42a434437964154

Observation 9f3b2e45-ec2a-407a-b6f0-eaddc4eb220f · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.130815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.130815Z digest=sha256:d54f5559855e7b1de13e08f7ec3afb961dda1c508fd20dfcaa4187a90a39a322

Observation 111207c7-2583-48df-907a-babeb134fe9c · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.114265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.114265Z digest=sha256:dcb60769e9de027f97e62285a67bc7a1cbf5105cf95b9e19ebe379b0ff5a871f

Observation cf8efb61-fa51-4edd-884f-b7ca1d56c4e0 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:03:12.256459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:ce3b84077837e3124ce852933a73fd571a0aff3ae7ea7f8d980569be9cd3db0b

Observation cfb4d057-5666-4f78-a646-a46b32205dc3 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:38:43.165665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:11:00.461167Z digest=sha256:0134f2e6443e9a6ad5087f6f0b41971538dd8929dd90553e1fae2f6a95584145

Observation d132e84b-6c28-4475-8e41-eb619ea4a887 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.869841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:b20bbd4382873aca1a509c0ed6cae9187763e345893e5553ae636dbcc6d35cfe

Observation ed168971-597e-4e1d-b4c1-736851e055d9 · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.576013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:e3b09d3cc8cff48ee394389a3a948df01f709289e5fee703d45125de96af63bb

Observation f45e8b3e-a9b3-4e73-9033-66b690564a28 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.359361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:06bebdcd4d2d1135fd927520dc3cb6e6a6f89386a46a89b77307c4d6aae0c3da