A handwritten bilingual VQA benchmark with about 1.5k pages and up to 4.8k question-answer pairs, plus baselines showing current models perform poorly, particularly with OCR-derived text.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 15 (CVPR) (2017) 1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
A handwritten bilingual VQA benchmark with about 1.5k pages and up to 4.8k question-answer pairs, plus baselines showing current models perform poorly, particularly with OCR-derived text.