Pith. sign in

Paper Citation Record · LEDGER

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

As of 10 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2502.07411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07411 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:54:56.474300Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20a166fa-0211-4747-8913-f525b65029fa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.114885Z digest=sha256:895eaf7b38736d4aa75a302d6d34df75d268305853a2fd92e7525cc7c04d0a5b

Observation 70b606bd-fcfd-4e71-ba3c-c5b514aa235a · outbound

This paper cites Where did i leave my keys? - episodic-memory-based question answering on egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Where did i leave my keys? - episodic-memory-based question answering on egocentric videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.120306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.120306Z digest=sha256:c0f7e2ee80a5141f6dc57faacc6b95c8e8ae2b068fe604adbd95b746ab6136bf

Observation 12c78ca8-94ce-4a70-bc4c-046db810186f · outbound

This paper cites Scene text visual question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Scene text visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.124737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.124737Z digest=sha256:1b376de4ceb8c04e1873100898384e348c5bbc40ba1a91c07002071a010f1b66

Observation 701d60ec-cf0d-4114-9199-9c1eea17b542 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.129066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.129066Z digest=sha256:f886d45a869c148115307acd634533f9fb9e5508f4b1c6fe7ed11a6945c929f9

Observation 6b7ebe36-a71a-44df-b74c-9fb4f39b937c · outbound

This paper cites InternLM2 Technical Report.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.133631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.133631Z digest=sha256:50edbcae366e9ade2247cfd10569b3579202ace20bd7175fdf21bde53ac5883e

Observation 6a3ad532-9edd-4f91-a7da-72196dccc641 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.138227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.138227Z digest=sha256:631d9608d5cc0355ba2a45c5135a073062a37947afdfcfddf9f8528a92279bc2

Observation adb75ff9-37fb-458b-a65c-3d2cbbcc3648 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.142716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.142716Z digest=sha256:1bd0d47def4f65873313809c8a00b0e164f5d8b22b52a9a2c7648121c432eadd

Observation d9661cac-6181-4b7a-bee4-f2a91c45d520 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.147036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.147036Z digest=sha256:b3c8a09d31fef2bbc9a9c5797f3d64d2d404ade459141d47c29117850a4462ef

Observation 4b68f2c0-dcc7-4251-8541-cfa9be868bbd · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.151522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.151522Z digest=sha256:ffb1394a315d6fd5910176ae4b68a79df706e244fd90766761c48523447ee811

Observation 5b98d8e6-29cf-4455-bb5c-51f67345d832 · outbound

This paper cites Egothink: Evalu- ating first-person perspective thinking capability of vision- language models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egothink: Evalu- ating first-person perspective thinking capability of vision- language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.156323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.156323Z digest=sha256:d942ff0fd19013324b578373096171bf4117c4e444eb29be06ae4864fa4f3be5

Observation cdef0dc8-631b-43ab-86e0-740dde11ec80 · outbound

This paper cites Grounded question-answering in long egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Grounded question-answering in long egocentric videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.160240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.160240Z digest=sha256:180abe187cd0430014d18aa42eed38525f777511d00901acd203226eea209e52

Observation 060f0cb2-c727-4ea5-a003-85cc2ce94246 · outbound

This paper cites Egovqa-an egocentric video question answer- ing benchmark dataset.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egovqa-an egocentric video question answer- ing benchmark dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.164622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.164622Z digest=sha256:819494b62d3c07d9da8ebc6fba43b7e446ee7da090aa722fcf74237cddc54982

Observation 320ac191-1522-4b26-8062-0dac98c0c302 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.168704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.168704Z digest=sha256:64d524074b5ee759055ec4ad953e9d4550126410781c8ce098f121d0d58cd440

Observation 4a1f8428-c2b7-4097-9d5c-e43294e33dfe · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ego4d: Around the world in 3,000 hours of egocentric video

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.173248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.173248Z digest=sha256:f48ea3094a48c193ec761500035e18866a4d86f1976069c21fee2b916ba36f0f

Observation cd64bf27-6376-4f56-853d-a0148fe9ab1b · outbound

This paper cites Context-aware graph inference with knowledge distillation for visual dialog.IEEE TPAMI, 44(10):6056–6073, 2021.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Context-aware graph inference with knowledge distillation for visual dialog.IEEE TPAMI, 44(10):6056–6073, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.611618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.177341Z digest=sha256:b519edfd00b9d92a182ae5a3305472c9326e0fcb130180c43b96928d1f3b3755

Observation 9cf3228b-03da-46f6-970d-c11f31eb9879 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Vizwiz grand challenge: Answering visual questions from blind people

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.598545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.181352Z digest=sha256:62afa439fd200087a84765c4c23a865bf7aa93a9390454a34f8e477628c5a944

Observation 9a8936b8-d246-4859-9ec9-d5c652b389c6 · outbound

This paper cites GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:54:56.771045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.185230Z digest=sha256:886c40f9498adb754ee0f3101a638a1599f1480e76669ae40fe3b679b36b0e7c

Observation b749407b-50fe-48b0-a978-ce8c3d8f9694 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM2: Visual Language Models for Image and Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.189601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.189601Z digest=sha256:db61ea0a98aab91a96ce5241f8180e6e556919ec02a178432aae755853797913

Observation a9b38a20-4d6d-4bd2-b78b-d9538c0a23b8 · outbound

This paper cites Understanding video scenes through text: Insights from text-based video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Understanding video scenes through text: Insights from text-based video question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.585323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.193845Z digest=sha256:411525f04c47b0db5d305b32afb4d0be3b3048b33a597a9a5be28dcf455bc1c3

Observation 1e950c31-aac4-469b-89d8-369b270623c6 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egotaskqa: Understanding human tasks in egocentric videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.571518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.197727Z digest=sha256:4b81e6b31f90bbfec1b97b7d9cda28711c9a60add7f7f3ff2b193a0de553b4f4

Observation de1cdeb5-a53e-4964-9007-ac5b3d5ba440 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.558677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.201725Z digest=sha256:eb9b36d0e1eed33ffc6f982ca559e4899e0eed95022b6219d7d176339a754101

Observation 7ddea211-36e8-41b0-a85b-d97653ced7d2 · outbound

This paper cites PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.205791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.205791Z digest=sha256:3ad6ad8c8c865affc5fab68444b85bf50578f1f7f71aaed5dab252748dbfc4a1

Observation bc3bc08a-889d-45db-ab84-ad27feba1338 · outbound

This paper cites Flex- attention for efficient high-resolution vision-language mod- els.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Flex- attention for efficient high-resolution vision-language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.546062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.209825Z digest=sha256:2d37034cc6d842176486796879d4ecd19ae578a187375de2a301ec88e8122b09

Observation e573457e-0ffe-4b46-9ccb-2873f6e870fd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.532537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.213744Z digest=sha256:fcc2b44acd7de073aa86cac0cfb83794488280cd648be93b41e4822e03e05952

Observation ad03128f-81c0-439f-b859-7bbc660cd8ab · outbound

This paper cites Invariant grounding for video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Invariant grounding for video question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.519362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.217734Z digest=sha256:ce30c0ea3a67d0bbc3ae2f52968f6d41196381ce4ca12ec85dbec6b86a37b6a7

Observation 11d2956c-034d-4817-8ebd-446f2ec62029 · outbound

This paper cites Transformer-empowered invariant grounding for video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Transformer-empowered invariant grounding for video question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.506144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.221976Z digest=sha256:9f4e74c2ac459461ae1d34435748b034f008e832d78481041b5a7b3dc07ba1db

Observation 832a7826-8ac0-4a40-9743-48d429db36a9 · outbound

This paper cites Vila: On pre-training for visual language models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Vila: On pre-training for visual language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.492342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.225880Z digest=sha256:20e2c1c1a7717098912260affb3ecb0aa0e0a80ca088d07e21566c48410dcbfc

Observation 35ec54d9-b5ae-410c-a685-7a899fc47138 · outbound

This paper cites Egocentric video-language pretraining.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egocentric video-language pretraining

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.478968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.229826Z digest=sha256:e97e125f9f8b6ff36028225addf828e7669b6e98cdd25caeff959e8d55f1dad2

Observation afa12dbe-bca5-4a39-8c19-f12ce12055ac · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.465382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.233949Z digest=sha256:335881c2b8094937551f4976de82343b4bf2c61bf6f7b806f00cefd54db909cf

Observation 3f9cd5da-9089-41e9-9dc0-b25fbb69f521 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.451768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.238033Z digest=sha256:6fe3f70caf285648f05c16415c4667fea6d8d79f37c66cfd6b37f83adbc535de

Observation 479d4a04-a8e8-44c7-89c6-c9552dd96ba4 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.241899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.241899Z digest=sha256:ffef80bbe385754e2fb5a8a916c7f927267edabfa5c4f9f4f5bd9f8b23f7c2cb

Observation 021e9583-38fd-4bca-9973-fa36d7b69e7d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.246341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.246341Z digest=sha256:2d9a3d7e489ab3f43f5c820bb41c9657c57147c1acd23e24853540589e6e751b

Observation 47d7ef9f-b74f-4650-85d7-479402dc0a49 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.437157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.250465Z digest=sha256:84477fb282b540c0dce5197391d281c2df5ed6dfa8a426d6034d900b219a17b7

Observation 6e4a0a78-fd94-4af5-9336-0400dd24e2b4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Docvqa: A dataset for vqa on document images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.423250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.254547Z digest=sha256:91e90decdce2cdc681c3f90e8c08eb35565b44b11db04b60da65c9b964950695

Observation 078de082-67ba-4d79-abcc-6bc671999581 · outbound

This paper cites Infographicvqa.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Infographicvqa

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.409491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.258520Z digest=sha256:9727f818b34e5f41038df7f74f11ae06486a2030bdcdd686034ad9ecd39c4079

Observation 78603dc3-316d-4193-9e7b-c32dfcbcb91e · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ocr-vqa: Visual question answering by reading text in images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.395858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.262488Z digest=sha256:f557c1e3f0d178e4a076fc7fc3d56795366d2f5eb2e4ccc9b1985295e19c2b51

Observation ee178bbd-26bf-4bf4-ab7a-c2d1842cf9e8 · outbound

This paper cites Gpt-4o system card.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gpt-4o system card

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.382022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.266298Z digest=sha256:25da76fb34d2db51a4be4b002bd41f167910a1c4d868aa74d5caa53d901d3dee

Observation d88a5540-81ff-4582-ba0e-db44de5c68fd · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gpt-4o mini: advancing cost-efficient intelligence

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.270226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.270226Z digest=sha256:a4b00043c3386acddf92ea330e6af4b3c80ee8d7473c5aff0607d641bc210ad2

Observation 2b8c4d99-595b-4aa8-b55d-507fefc2139a · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.359826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.274244Z digest=sha256:a5e3caff54c1695a2976a0c973ed4589bcd217bdf4a4bed7ea56c61c394a0242

Observation 1a549ecd-6650-4be7-97ab-26169dfee7ff · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Learn- ing transferable visual models from natural language super- vision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.346472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.278216Z digest=sha256:5cd1cea5330a5ce7e707865e8f88040f63e2c4f1814ca4460f95b681f2bb3bdf

Observation b40979d0-1bb6-450f-af94-f324be730f5a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.282084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.282084Z digest=sha256:a202a70f539a5ea5978db83784f5417aa818a408f993c449ba2359404344bc9e

Observation 79a9a262-70b9-4325-b797-1de2079f524e · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.286985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.286985Z digest=sha256:3562d667ad2be4b4fb42e14744620fdef617c80319affdcfb26ba6979d2edc7c

Observation bd87f972-824e-4eb2-943b-4e0dad9ca5a9 · outbound

This paper cites Annotating objects and relations in user- generated videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Annotating objects and relations in user- generated videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.333415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.291155Z digest=sha256:7c004205c9d78ac9011ea3c9273a350c646142be279eb12490c8a88f55e61ad5

Observation ad865063-2f7f-46ff-8da0-c522eb334193 · outbound

This paper cites GLU Variants Improve Transformer.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering GLU Variants Improve Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.295240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.295240Z digest=sha256:6a0e9969dc41a0bab0b4b2effccee2e6227134aeef591231a0e199d673455b96

Observation 1f35c8e7-ee3a-4a8a-b152-e8ba7621aba7 · outbound

This paper cites Towards vqa models that can read.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Towards vqa models that can read

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.319919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.299537Z digest=sha256:ee2130323d5c0c7e009031baff9c642df5f8c9a26a0ade39e3cfb04b1e07a95b

Observation 7f75d1ff-79d6-49ef-ad9d-e34462632914 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.303789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.303789Z digest=sha256:c494d5319faced1b76e1cb454d14246ca052b2f58d5bd66828b643873120fea3

Observation 322010f4-f1b3-4cb8-b534-b2a38c6f2186 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.308154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.308154Z digest=sha256:dafed3c5f4172fcb5a7010ae8c4838e1feb1888ef9f374f7f7a2f3f6d3a3f883

Observation 1c5febe1-94b6-4db5-9200-00f1c50ca31b · outbound

This paper cites Reading between the lanes: Text videoqa on the road.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Reading between the lanes: Text videoqa on the road

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.306916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.312592Z digest=sha256:f32462c149e9e906409901638e91bb8b53658b878d9ae4c0faec11e538554e01

Observation ac6caca4-760a-4449-ae7f-b41e840b3571 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.317080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.317080Z digest=sha256:d8f7d4df43b51f0fcaad1c446e6ac203f4f0e1cde223e2c9a7d3d0411fb41fe9

Observation 505a1408-9d54-4ed5-a34d-ef627ddcb886 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.321415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.321415Z digest=sha256:6acfe3be35b1fcae4b22a87e042f7554d4c7c4216777d100f3048ba24dfcf95e

Observation e0dce2cd-fd45-4e5c-8037-27e7602e7d54 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.325636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.325636Z digest=sha256:c5311ab4b0d8d42eeb3b384c002360414fe104516898ce84d2f394007f576a61

Observation ce44bf7e-1841-465e-b4d6-66bba7d0d2bb · outbound

This paper cites On the general value of ev- idence, and bilingual scene-text visual question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering On the general value of ev- idence, and bilingual scene-text visual question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.293858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.329931Z digest=sha256:7096452b2cea00e91b67ab35ef7d221e12b6fbd78d35fff45d6f64a7d0895ca9

Observation 3c6960a3-522f-49e1-924d-5beab171956a · outbound

This paper cites Assistq: Affordance-centric question-driven task completion for ego- centric assistant.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Assistq: Affordance-centric question-driven task completion for ego- centric assistant

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.280040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.334020Z digest=sha256:a2d56ac8d57475304c79f871ac8241e5f6ea1f6a2a91b533a233572f8f7d7474

Observation 21504019-c8a1-4b2a-b024-8bcaaa3c39ae · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.266470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.337940Z digest=sha256:4b6209fa7ff11df2f7926a45e46428faa5776108f0c68cd5f5c5c8e3571ce036

Observation bc512d6f-164a-4ff4-91ac-0044c19ddbab · outbound

This paper cites VideoQA in the Era of LLMs: An Empirical Study.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VideoQA in the Era of LLMs: An Empirical Study

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.342107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.342107Z digest=sha256:da7e5f290f394475ec9f83dca56200a65ba8b023668dd3bbdf056cc5b37ba5f7

Observation de7ca57f-0f03-4635-ba80-185cf6a168b2 · outbound

This paper cites Deconfounded video moment retrieval with causal intervention.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Deconfounded video moment retrieval with causal intervention

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.252848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.346273Z digest=sha256:6cd5716301fd84ac61e6fb13e71d974876c6820c9770ba02d6debdba7c769e32

Observation 579d69eb-bbc1-4f7b-b01e-272581b454eb · outbound

This paper cites Video moment retrieval with cross-modal neural architecture search.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video moment retrieval with cross-modal neural architecture search

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.239274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.350327Z digest=sha256:8281204d41aaba7d4c789dd03b60ec21f4bcd9a5f3f9b1519522e4b13447dd07

Observation dcb06e1c-a964-4ef3-bab3-768a016de2e9 · outbound

This paper cites Robust video question answer- ing via contrastive cross-modality representation learning.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Robust video question answer- ing via contrastive cross-modality representation learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.225574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.355088Z digest=sha256:800bb48c3e05766eae98b37eae5f6d226b897c63c5bc4e3259d89fc7107b526f

Observation 028fa900-a965-4130-b355-e1a364996573 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.359154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.359154Z digest=sha256:2a246d5b802f7797c316f3f45df9d93ee2791f4a6ae239e0079f8751ed3e5c02

Observation eb601c86-bfa8-4d66-b3ba-f2c41546f759 · outbound

This paper cites MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.363353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.363353Z digest=sha256:b3dacd007e59dd662af5f38dcd40f274de9fdb035fe1ec2ac484b0b28125cf68

Observation 9437a4c9-1fae-4824-a387-cdecb0997592 · outbound

This paper cites Sigmoid loss for language image pre-training.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.211122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.368065Z digest=sha256:2a781f33327678f8df1179d0d78011f3b882e7bf6b231a005b6886df2e54d7cb

Observation 743af68e-86a0-4e98-a5c4-6662b4f6faf3 · outbound

This paper cites Multi-factor adaptive vision selec- tion for egocentric video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Multi-factor adaptive vision selec- tion for egocentric video question answering

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.197499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.372009Z digest=sha256:c47b20bda5911290372edb28279edf1d7b74f59de42051702837d0640f38ead7

Observation 03acbac3-de00-4c32-926f-a2c60e18e971 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.375610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.375610Z digest=sha256:f07419d59100fb8c403243c9c6efb082049db449f3dacc3272dff92cf45017ab

Observation c964b9ad-5945-4630-94c1-15f72c0224b5 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava- next: A strong zero-shot video understanding model, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.183592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.379684Z digest=sha256:a3551f63b189881605f5ebe3bcdaedd5f5e09c2a54dc627c0483af8c7007d203

Observation 6490b45e-29a4-45bb-a171-635ebdb7bd6c · outbound

This paper cites Diffusion-based blind text image super-resolution.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Diffusion-based blind text image super-resolution

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.169518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.383600Z digest=sha256:0bc41e4c1ea4e469fc064c89d1e61366a511053e93b926f4593ac9964ef73a13

Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.387608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.387608Z digest=sha256:499d977469acedf72a4b57b9fb828103b73f64e028b207b80c1260391648b79b

Observation 0dd917b1-93d0-4ce9-94f5-f25fc4a8ae42 · outbound

This paper cites Towards video text visual question answering: Benchmark and baseline.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Towards video text visual question answering: Benchmark and baseline

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.156198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.391728Z digest=sha256:946c89336014a519c309d931d490bdbddad04dde040d88c7d9fb3922f0060e99

Observation 72714bec-23ec-4e34-a8fa-b3ab4e74f246 · outbound

This paper cites Exploring sparse spatial relation in graph inference for text- based vqa.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Exploring sparse spatial relation in graph inference for text- based vqa

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.142122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.395299Z digest=sha256:a515fee849c3f4fc0204e5143b62c6539fbac262097c3e3abb8d2ceae3302ee4

Observation 68071ac0-de17-4ee3-9072-9ca247484a51 · outbound

This paper cites Scene-Text Grounding for Text-Based Video Question Answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Scene-Text Grounding for Text-Based Video Question Answering

Reference 69

Resolution
malformed identifier
local_arxiv, observed 2026-08-08T12:54:56.516262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.399357Z digest=sha256:10ff19e999070cb49bb16c13e951c60c8803c6ba7872548e6ba34533d860505f

Observation 5cc3978e-f96d-4cba-a5aa-c51c18902e0e · outbound

This paper cites I” should be used appropriately. Requirement 5: The questions should be of moderate length. When announcing the question please label each question as “Question 1, 2, 3: {question}.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering I” should be used appropriately. Requirement 5: The questions should be of moderate length. When announcing the question please label each question as “Question 1, 2, 3: {question}

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.128528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.404806Z digest=sha256:53500e5eaa50298b4c77055aaf6679485041b3f2372f17f87b95a0e2a12c0fc3

Observation e2e47122-0bb7-4a7c-b3bd-dfa58fb80fa8 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.114701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.408988Z digest=sha256:6b89aa1f6196ae54c2c7838f8b0da978abffd94d21c4d6a09207c63ec58d1625

Observation 2b9872bd-9eab-43fe-9ccb-b8c93b47c37c · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.100405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.413149Z digest=sha256:ebbeeec4a816631e6acd4124a8030f223f72cf6a92d6da9d70321b7cdeec0764

Observation 458701d6-fef6-4b9e-9a3f-844702db5b7d · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.086757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.417105Z digest=sha256:08775c36b489470cb671453b9b3f647c5524b8421b1febf809a133992062b00f

Observation 54dec6c6-efc5-4252-97c2-4b18cbce608f · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.073350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.421837Z digest=sha256:8ea9e446b61e2488fe7084956fbefd40adfee00fb91e1a4e33b378a478327b81

Observation 82e9b652-e3b4-40f4-aa59-b4cecc0acba3 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.059763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.425714Z digest=sha256:c63602198c08a62b247c5c1ca0c0ab7a8ee780412e6b9cbf44d452feaa3ba5ab

Observation a1ed87b2-e23a-473e-a0a2-10f9877f776c · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.046256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.429588Z digest=sha256:8eedefe672dc8f80fc6e58aa2a0f4ff02ea1861b53b6472f55e7c3c48dca03e6

Observation a97fc98a-d790-4ae9-96b9-b64303857f51 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.032389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.433279Z digest=sha256:9b872eec4f98091595010366788976630c98fbc32ea5b9189113f9f062c778c0

Observation 0a2249fe-e1cd-4213-91f8-40bdbd6c6474 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.018820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.437721Z digest=sha256:453a9ebefce55d8f5696ce89d098c95c80aadbdad9ce9cba1509d2805289501a

Observation 5b4ca62c-4072-42f2-acf9-3f8579a0c459 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.004118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.442124Z digest=sha256:26f7b87b830736d399025411cb6bef9dac37ac9453d21d14278a150d8ab81e51

Observation 22df1ca5-ea0d-4615-beff-1780dbe5dafa · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.989360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.445784Z digest=sha256:d8cf40f1df94a48e7b3f0111fc7888986a6b7122b6ebbf67dc25a649c855a202

Observation c703cccc-3d52-47c6-aa33-797bd13546f4 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.975681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.449378Z digest=sha256:2f85e260966d5d6fd21e2e2e540ef5ceee2b8781b3c2246a555f4d5e7898baf4

Observation 9c8d3008-3003-4077-b3ff-77a4a12865ad · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.961082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.453801Z digest=sha256:bd4cebba8ec131b241875fe6e81204cfb71dc9cae6dd46e35d0121f00654f420

Observation 86832389-b959-4be6-add2-26981f4cbd4b · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.945488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.458157Z digest=sha256:8043bf96bfe81681a42efcd87035f52b54dae2929458d6c01cd0c0b18b59dde3

Observation 34ead312-430b-4a56-b89c-d83a5f1eb0a1 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.930849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.462137Z digest=sha256:59cfd68380730e9d1e8f5e074154870a2521f346badc9f3fbb7281d0faac7b72

Observation 6a77d3a1-a090-4eed-89fd-57360faa4222 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.916036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.465973Z digest=sha256:27fdf2c86f1d6147d350b95352ba0036a4bfe4b8251bed95cefb81bb623be24f

Observation 13fb9799-69ef-463f-8e69-265da5d13af4 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.901345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.470475Z digest=sha256:b4dae9aa49469c8f0faa96eab942e2150109c46467bf1e9434b242544a54df1a

Observation c69b76a0-2013-4a16-bcb6-6290905136b1 · outbound

This paper cites Unanswerable.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unanswerable

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.886834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T12:54:56.474300Z digest=sha256:8da1cebc1b2afee4e48d090887cddd365cba4fa13085f6fdd3f6fb52270f871a

Pith citing papers

No inbound Pith citation observations are available.