Pith. sign in

Paper Citation Record · LEDGER

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2502.07411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07411 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:54:56.474300Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20a166fa-0211-4747-8913-f525b65029fa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.114885Z digest=sha256:78f5566a07fa6f00a8b9493f65c955e5acf9796cee64151a9769eb60da5782d8

Observation 70b606bd-fcfd-4e71-ba3c-c5b514aa235a · outbound

This paper cites Where did i leave my keys? - episodic-memory-based question answering on egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Where did i leave my keys? - episodic-memory-based question answering on egocentric videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.120306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.120306Z digest=sha256:5a12a82a404c9181a2e298916a4219624609541ee3e11617a641d67b34cce74e

Observation 12c78ca8-94ce-4a70-bc4c-046db810186f · outbound

This paper cites Scene text visual question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Scene text visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.124737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.124737Z digest=sha256:cd8f66eb8982596c9519182ecd0d82f5672b0443b5415f7354b38047c539871d

Observation 701d60ec-cf0d-4114-9199-9c1eea17b542 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.129066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.129066Z digest=sha256:2dc8aa3636310100663acc2f4fd2ddcb5d65a3bab6b01381910956f59547996a

Observation 6b7ebe36-a71a-44df-b74c-9fb4f39b937c · outbound

This paper cites InternLM2 Technical Report.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.133631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.133631Z digest=sha256:d788595bc630f9b56e22efd526d7d4a0b11d47770301f6dbe06a3d1b7b136b18

Observation 6a3ad532-9edd-4f91-a7da-72196dccc641 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.138227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.138227Z digest=sha256:ae14cb966b0a6f339ca56a0b1d74cb03f846f1598af8e2d83e455789d949e177

Observation adb75ff9-37fb-458b-a65c-3d2cbbcc3648 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.142716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.142716Z digest=sha256:c211029ca273632d269b4915d0c3a53841d23fab9b269ecb49cbf7e1ba696852

Observation d9661cac-6181-4b7a-bee4-f2a91c45d520 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.147036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.147036Z digest=sha256:a1821a20c5145f239444076180aaf8d94ce5d5fa6be21673db4a0a9999502eb7

Observation 4b68f2c0-dcc7-4251-8541-cfa9be868bbd · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.151522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.151522Z digest=sha256:f9099a7a99f360d3e8e2c8973ab21c9e5b0cc7fa98dfa70030b53c4c34d61ec1

Observation 5b98d8e6-29cf-4455-bb5c-51f67345d832 · outbound

This paper cites Egothink: Evalu- ating first-person perspective thinking capability of vision- language models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egothink: Evalu- ating first-person perspective thinking capability of vision- language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.156323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.156323Z digest=sha256:f7cb65dbd1465c02e4db06bdaa032d823dda922a7f45a2cda41c0cd6c61124d1

Observation cdef0dc8-631b-43ab-86e0-740dde11ec80 · outbound

This paper cites Grounded question-answering in long egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Grounded question-answering in long egocentric videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.160240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.160240Z digest=sha256:aa3a9f50a39c0c321070881c4ec9dbc6300c81d7327e54f2729efc5206dc9ce7

Observation 060f0cb2-c727-4ea5-a003-85cc2ce94246 · outbound

This paper cites Egovqa-an egocentric video question answer- ing benchmark dataset.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egovqa-an egocentric video question answer- ing benchmark dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.164622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.164622Z digest=sha256:04e80165a16cd11a8b950cc794545a410a728cd70a24ec7bcd9c8b496488efc9

Observation 320ac191-1522-4b26-8062-0dac98c0c302 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.168704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.168704Z digest=sha256:659b261cc951bbd1c3b99807d5ad4821c8509e425d53a97f09f252da75b2ce3b

Observation 4a1f8428-c2b7-4097-9d5c-e43294e33dfe · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ego4d: Around the world in 3,000 hours of egocentric video

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.173248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.173248Z digest=sha256:e994e3650a3710e2ad8cc3035ec84a9431d1caa73cc5b68ae86b2c212ac35b51

Observation cd64bf27-6376-4f56-853d-a0148fe9ab1b · outbound

This paper cites Context-aware graph inference with knowledge distillation for visual dialog.IEEE TPAMI, 44(10):6056–6073, 2021.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Context-aware graph inference with knowledge distillation for visual dialog.IEEE TPAMI, 44(10):6056–6073, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.611618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.177341Z digest=sha256:2c6dbc3fbc4eb08950f6ee9a11d4954ecfb7dd9a41334d868c39d8964c2f9a37

Observation 9cf3228b-03da-46f6-970d-c11f31eb9879 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Vizwiz grand challenge: Answering visual questions from blind people

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.598545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.181352Z digest=sha256:8068e1249ce62b2b87a0f04ac5e21f2f7595d059b3c04724818a77116a0d591d

Observation 9a8936b8-d246-4859-9ec9-d5c652b389c6 · outbound

This paper cites GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:54:56.771045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.185230Z digest=sha256:8938b6a0c68398068b680c33b0fc62ff4543b5f8f06ed0eeeaa6f5f219e5b561

Observation b749407b-50fe-48b0-a978-ce8c3d8f9694 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM2: Visual Language Models for Image and Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.189601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.189601Z digest=sha256:1580530c91fc300a138102142e3b3aea7e5915c9d88767cf2b2c49a1ecd5e484

Observation a9b38a20-4d6d-4bd2-b78b-d9538c0a23b8 · outbound

This paper cites Understanding video scenes through text: Insights from text-based video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Understanding video scenes through text: Insights from text-based video question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.585323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.193845Z digest=sha256:e7674c8eda554e3af82752da05b38db922814489d55e88089340510baf2e0e8a

Observation 1e950c31-aac4-469b-89d8-369b270623c6 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egotaskqa: Understanding human tasks in egocentric videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.571518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.197727Z digest=sha256:dea38988b7be0ce51ac18adad60358a7e4a0843a906375784ed7d913ef97c724

Observation de1cdeb5-a53e-4964-9007-ac5b3d5ba440 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.558677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.201725Z digest=sha256:1b8317d30812f5612109dc70829121e0f38a93e7bd93d440051e864b327a85db

Observation 7ddea211-36e8-41b0-a85b-d97653ced7d2 · outbound

This paper cites PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.205791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.205791Z digest=sha256:d19271a763234e769b0df14bb7751450f97fbe2d14b54e4d7ee5fd908b0b0009

Observation bc3bc08a-889d-45db-ab84-ad27feba1338 · outbound

This paper cites Flex- attention for efficient high-resolution vision-language mod- els.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Flex- attention for efficient high-resolution vision-language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.546062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.209825Z digest=sha256:6b3636c89190bcc5a4e83db4c99b08ec6e46e8d3a7e71c3c21f838a5a6e2b071

Observation e573457e-0ffe-4b46-9ccb-2873f6e870fd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.532537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.213744Z digest=sha256:b4069734ce28e07c2a63bbd6ad19e57547f0a44b06938f391bf37a41859e3d47

Observation ad03128f-81c0-439f-b859-7bbc660cd8ab · outbound

This paper cites Invariant grounding for video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Invariant grounding for video question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.519362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.217734Z digest=sha256:1d3fc7c8347d65b651cadd067f901dae8ffbf82258fddf1e2c2fd72b9c436534

Observation 11d2956c-034d-4817-8ebd-446f2ec62029 · outbound

This paper cites Transformer-empowered invariant grounding for video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Transformer-empowered invariant grounding for video question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.506144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.221976Z digest=sha256:3fb4767f78137173d68ba85501184d26a80b8d1903447294fb11295f8d1bfe56

Observation 832a7826-8ac0-4a40-9743-48d429db36a9 · outbound

This paper cites Vila: On pre-training for visual language models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Vila: On pre-training for visual language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.492342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.225880Z digest=sha256:d09e6f1454709902a9a27d30d9b080287e24675d590b75611a18640eda2597bc

Observation 35ec54d9-b5ae-410c-a685-7a899fc47138 · outbound

This paper cites Egocentric video-language pretraining.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egocentric video-language pretraining

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.478968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.229826Z digest=sha256:98f78e2cd2fbd0b53eb2b49f1a58fe066406e6f9f7fa6cf1d38c4b1159f58775

Observation afa12dbe-bca5-4a39-8c19-f12ce12055ac · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.465382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.233949Z digest=sha256:044ed31237565d1bb48cfd2fa9ba9f900507d085af7fcc5eaec28e2321611248

Observation 3f9cd5da-9089-41e9-9dc0-b25fbb69f521 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.451768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.238033Z digest=sha256:98d70a4e34f62601b48f484cc5c95149f1fb771fa14be65b2c2362ceb4109633

Observation 479d4a04-a8e8-44c7-89c6-c9552dd96ba4 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.241899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.241899Z digest=sha256:175bb3299a6941fe39767a037f1f3ebb91bebe8684934b3ff999a3ec2e6889cd

Observation 021e9583-38fd-4bca-9973-fa36d7b69e7d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.246341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.246341Z digest=sha256:0e505864c4231b7870536ec31d068122842a87f52dd389d63aa9fcd2abe5098b

Observation 47d7ef9f-b74f-4650-85d7-479402dc0a49 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.437157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.250465Z digest=sha256:52243008703d01147bdbf0d30c9a6acf0a47f7d5ed1300b51d50d5d361351b39

Observation 6e4a0a78-fd94-4af5-9336-0400dd24e2b4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Docvqa: A dataset for vqa on document images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.423250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.254547Z digest=sha256:44593b8353d051daa156745852e24d08bbd5f2466942b12927c34b55a3d6aee7

Observation 078de082-67ba-4d79-abcc-6bc671999581 · outbound

This paper cites Infographicvqa.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Infographicvqa

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.409491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.258520Z digest=sha256:e030b4fe113c7d048fd7db4811541e557c5856ed9fb2750174a8cef412d616d8

Observation 78603dc3-316d-4193-9e7b-c32dfcbcb91e · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ocr-vqa: Visual question answering by reading text in images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.395858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.262488Z digest=sha256:7ff2c04e3066398f57f287d4ebf5f8306d02cdb43250397e607e6ff8ddfe7636

Observation ee178bbd-26bf-4bf4-ab7a-c2d1842cf9e8 · outbound

This paper cites Gpt-4o system card.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gpt-4o system card

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.382022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.266298Z digest=sha256:8c403a9ff5975b3bb777651a75e7d6c0bf1f21a37ed4101f4c298750aabba869

Observation d88a5540-81ff-4582-ba0e-db44de5c68fd · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gpt-4o mini: advancing cost-efficient intelligence

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.270226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.270226Z digest=sha256:5c45cb8e938c854e18f3203d6f72288b72b1093cdcf7660e67440ba7eca16492

Observation 2b8c4d99-595b-4aa8-b55d-507fefc2139a · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.359826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.274244Z digest=sha256:a5f757673dfa9daf5289a74d68f84bda297ce5fb49005be39ffb13c8f70b5010

Observation 1a549ecd-6650-4be7-97ab-26169dfee7ff · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Learn- ing transferable visual models from natural language super- vision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.346472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.278216Z digest=sha256:4484c20342c83d5d9d94e85641f58607e8348f4cb27d954802a8ede54046c415

Observation b40979d0-1bb6-450f-af94-f324be730f5a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.282084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.282084Z digest=sha256:a59dc345fa194d0a38277201881ae89df894dfbb70131089cc036c54aaff7083

Observation 79a9a262-70b9-4325-b797-1de2079f524e · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.286985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.286985Z digest=sha256:cfb5061a7e8ccd612b7d10e329ce5f61cc62bf6734394a22e5207aa8c03b0c17

Observation bd87f972-824e-4eb2-943b-4e0dad9ca5a9 · outbound

This paper cites Annotating objects and relations in user- generated videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Annotating objects and relations in user- generated videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.333415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.291155Z digest=sha256:e3558af627f4d74e0be9bbe9ead47d6689a5394fb359bc3d0724d4b713868bb1

Observation ad865063-2f7f-46ff-8da0-c522eb334193 · outbound

This paper cites GLU Variants Improve Transformer.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering GLU Variants Improve Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.295240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.295240Z digest=sha256:88bdae93d376a68dc10dddd2aa42670297aedba7b135cd507f4380d8a74ed967

Observation 1f35c8e7-ee3a-4a8a-b152-e8ba7621aba7 · outbound

This paper cites Towards vqa models that can read.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Towards vqa models that can read

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.319919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.299537Z digest=sha256:854cc85a9b53166e73ea2a1b9f4897d67227a719607bf86eda8c587d5d72baba

Observation 7f75d1ff-79d6-49ef-ad9d-e34462632914 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.303789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.303789Z digest=sha256:0e9af69210d2ec554df5b39137bd9503cf15f011b80eecddb9edc345e6c9c515

Observation 322010f4-f1b3-4cb8-b534-b2a38c6f2186 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.308154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.308154Z digest=sha256:29052fe3709829e4a42c5597c5161b109e64188af751d48efbbe0159d8bc7201

Observation 1c5febe1-94b6-4db5-9200-00f1c50ca31b · outbound

This paper cites Reading between the lanes: Text videoqa on the road.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Reading between the lanes: Text videoqa on the road

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.306916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.312592Z digest=sha256:50006d7ed9d68d77324d07a4e7f3d6e68e2482bdb53a8b55dd224a037b7fcf81

Observation ac6caca4-760a-4449-ae7f-b41e840b3571 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.317080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.317080Z digest=sha256:bf0becebd34344636a12859b911752c540fcec55c9acbd61ee956e9f2a42c426

Observation 505a1408-9d54-4ed5-a34d-ef627ddcb886 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.321415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.321415Z digest=sha256:687882d482f540d42db6d577305c0e92c7e3c2cbf77e95c9e0a85e0b6afa3349

Observation e0dce2cd-fd45-4e5c-8037-27e7602e7d54 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.325636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.325636Z digest=sha256:b4cf283bbcf90f011552c4e44576688cad4b2498184b2b7260791d7263c9d565

Observation ce44bf7e-1841-465e-b4d6-66bba7d0d2bb · outbound

This paper cites On the general value of ev- idence, and bilingual scene-text visual question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering On the general value of ev- idence, and bilingual scene-text visual question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.293858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.329931Z digest=sha256:ca417f4ac85a65b6213fc6154dd3c281bdc4cd6ee77a03236a39fb3574a5b26b

Observation 3c6960a3-522f-49e1-924d-5beab171956a · outbound

This paper cites Assistq: Affordance-centric question-driven task completion for ego- centric assistant.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Assistq: Affordance-centric question-driven task completion for ego- centric assistant

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.280040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.334020Z digest=sha256:b1a6d108ae3d74d409b1e216373ff0d56745d2f392e93690e850898ee3b08e19

Observation 21504019-c8a1-4b2a-b024-8bcaaa3c39ae · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.266470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.337940Z digest=sha256:280d9df89dfa95175c6d25bcc7e1fe7bb8840dc4920b31eaf8691f88c1a34a9e

Observation bc512d6f-164a-4ff4-91ac-0044c19ddbab · outbound

This paper cites VideoQA in the Era of LLMs: An Empirical Study.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VideoQA in the Era of LLMs: An Empirical Study

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.342107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.342107Z digest=sha256:383df77158bfb1ab4cd2ac8d13f166f2e538e0a8a982f9734a71910e98928dcb

Observation de7ca57f-0f03-4635-ba80-185cf6a168b2 · outbound

This paper cites Deconfounded video moment retrieval with causal intervention.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Deconfounded video moment retrieval with causal intervention

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.252848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.346273Z digest=sha256:33ed05c9b5df20938327f3b5e7bfd2d8a6004f7b21574313306ff1136ce0fc5c

Observation 579d69eb-bbc1-4f7b-b01e-272581b454eb · outbound

This paper cites Video moment retrieval with cross-modal neural architecture search.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video moment retrieval with cross-modal neural architecture search

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.239274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.350327Z digest=sha256:f6aee16e8b6fb599aacf3db592ed975bace041fd22864d8311a10c81ece40705

Observation dcb06e1c-a964-4ef3-bab3-768a016de2e9 · outbound

This paper cites Robust video question answer- ing via contrastive cross-modality representation learning.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Robust video question answer- ing via contrastive cross-modality representation learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.225574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.355088Z digest=sha256:d3b934034b633d36c11d5375a639e5f43ea2c144cc735f22817de12239dbf31a

Observation 028fa900-a965-4130-b355-e1a364996573 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.359154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.359154Z digest=sha256:67f80049fa8dcb5b356a02e75b619c7d9e31e751d4548824763c507f3017724e

Observation eb601c86-bfa8-4d66-b3ba-f2c41546f759 · outbound

This paper cites MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.363353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.363353Z digest=sha256:97dc4d769c55cbf83ab781186cb52dc22c31c305f177787a2a3164b439fa3e0d

Observation 9437a4c9-1fae-4824-a387-cdecb0997592 · outbound

This paper cites Sigmoid loss for language image pre-training.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.211122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.368065Z digest=sha256:583189bdb8ff9e180dfd2f4b66e4733fd20cb28e35bab542030bef2381ddea78

Observation 743af68e-86a0-4e98-a5c4-6662b4f6faf3 · outbound

This paper cites Multi-factor adaptive vision selec- tion for egocentric video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Multi-factor adaptive vision selec- tion for egocentric video question answering

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.197499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.372009Z digest=sha256:27b054d4534ce1e77dd0516ebbc50f946e95234fe28d2892b31dac65f47f3078

Observation 03acbac3-de00-4c32-926f-a2c60e18e971 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.375610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.375610Z digest=sha256:f0c28f194ac335dd5156ac51e24051320cb71ed809ef07b6596703ef1de23feb

Observation c964b9ad-5945-4630-94c1-15f72c0224b5 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava- next: A strong zero-shot video understanding model, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.183592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.379684Z digest=sha256:936bba921d4d2c7758937ac9b7bf66812d5e44242dc6a9eaf8d04c903f900f0b

Observation 6490b45e-29a4-45bb-a171-635ebdb7bd6c · outbound

This paper cites Diffusion-based blind text image super-resolution.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Diffusion-based blind text image super-resolution

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.169518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.383600Z digest=sha256:f4625b48e34637a445595b545f756b264e61543703202f08775165c9cdfe700b

Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.387608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.387608Z digest=sha256:931a21c22c17c9967d39c8bf9ef23c74821108ceeae663bb7bdef19d7bd57b72

Observation 0dd917b1-93d0-4ce9-94f5-f25fc4a8ae42 · outbound

This paper cites Towards video text visual question answering: Benchmark and baseline.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Towards video text visual question answering: Benchmark and baseline

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.156198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.391728Z digest=sha256:4b8d0c29733783549c02d78659c61d7f2bd8904791cf643d09051bca6f069ba7

Observation 72714bec-23ec-4e34-a8fa-b3ab4e74f246 · outbound

This paper cites Exploring sparse spatial relation in graph inference for text- based vqa.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Exploring sparse spatial relation in graph inference for text- based vqa

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.142122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.395299Z digest=sha256:baa07aa2794c695c1eda703e5fe0f5ca9ca5c13ad874e3b40b250aa5e6547a6a

Observation 68071ac0-de17-4ee3-9072-9ca247484a51 · outbound

This paper cites Scene-Text Grounding for Text-Based Video Question Answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Scene-Text Grounding for Text-Based Video Question Answering

Reference 69

Resolution
malformed identifier
local_arxiv, observed 2026-08-08T12:54:56.516262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.399357Z digest=sha256:9d7fc41d668cbceb1550010fa1ba6d63219bf9f35073d01c189eeba1c3884b87

Observation 5cc3978e-f96d-4cba-a5aa-c51c18902e0e · outbound

This paper cites I” should be used appropriately. Requirement 5: The questions should be of moderate length. When announcing the question please label each question as “Question 1, 2, 3: {question}.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering I” should be used appropriately. Requirement 5: The questions should be of moderate length. When announcing the question please label each question as “Question 1, 2, 3: {question}

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.128528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.404806Z digest=sha256:bc5c6e549ebb5710c5a919ed39f0553e60dee43b165382a0701927a3f9f5e649

Observation e2e47122-0bb7-4a7c-b3bd-dfa58fb80fa8 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.114701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.408988Z digest=sha256:86793db4a21dfc631e708be797ed5fa817a513eb53a0d3ea40742446589304ee

Observation 2b9872bd-9eab-43fe-9ccb-b8c93b47c37c · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.100405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.413149Z digest=sha256:827a001140b6a7acb80e68d1917e01401678cd19ed39b7b093732f649892c640

Observation 458701d6-fef6-4b9e-9a3f-844702db5b7d · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.086757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.417105Z digest=sha256:ec645084b8186dfca7d90908eddf3681036e2412d8836da25c05da2770b94c9c

Observation 54dec6c6-efc5-4252-97c2-4b18cbce608f · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.073350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.421837Z digest=sha256:18fba37700545a800df8d90f0d879c902cdc7ec66ed9895a8dbb512a0aac192d

Observation 82e9b652-e3b4-40f4-aa59-b4cecc0acba3 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.059763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.425714Z digest=sha256:66967e6f0f8ec06fd4a75bd77d8e3eeb824d38919ed02b9ac194f43229619018

Observation a1ed87b2-e23a-473e-a0a2-10f9877f776c · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.046256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.429588Z digest=sha256:f1c414a0d2113a362f38e626380ec456e86c95a46737ad4c2679a34c130d9310

Observation a97fc98a-d790-4ae9-96b9-b64303857f51 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.032389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.433279Z digest=sha256:bf2afcf594783a8bae53072d339dcbbbf6928d63f905b24821645f626ec62a9c

Observation 0a2249fe-e1cd-4213-91f8-40bdbd6c6474 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.018820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.437721Z digest=sha256:5813bb95f163e11170cb74530da764f8a1da3a7ffa4ba1ca1ed1ec57b479c6cd

Observation 5b4ca62c-4072-42f2-acf9-3f8579a0c459 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.004118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.442124Z digest=sha256:ed7c7d5635adac172ad59d5327c055da2a0ce964f81b325d98b5a6de64d175f3

Observation 22df1ca5-ea0d-4615-beff-1780dbe5dafa · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.989360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.445784Z digest=sha256:a565c788766ac0b2adb485bc28c41feada4ba403bdcc07cb3790c0c823047c84

Observation c703cccc-3d52-47c6-aa33-797bd13546f4 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.975681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.449378Z digest=sha256:7776221f71279e86cf085c2410d06b2cb9e2aa67c8c33306cea4a43e8872a2c1

Observation 9c8d3008-3003-4077-b3ff-77a4a12865ad · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.961082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.453801Z digest=sha256:5e373b51708529f184addfc94c8e0ef7e18900460695ea0d6f17d48693380a35

Observation 86832389-b959-4be6-add2-26981f4cbd4b · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.945488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.458157Z digest=sha256:6ba2caca12d9956bb4c1aa751724695a97e3f61469c41d78a7e2acef8b2fe294

Observation 34ead312-430b-4a56-b89c-d83a5f1eb0a1 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.930849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.462137Z digest=sha256:02433caaaabdf8da089c6185b84dac7946e2da4ef46b492b43fd5920d86ae6bd

Observation 6a77d3a1-a090-4eed-89fd-57360faa4222 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.916036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.465973Z digest=sha256:9bbdc2dd376d18fb44f7e3c1a3fe5809bfce2c002c6e27421ce373f479583b15

Observation 13fb9799-69ef-463f-8e69-265da5d13af4 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.901345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.470475Z digest=sha256:c55aa24c70dbaa5a490b5900dccf569a600ba4296dc25dfbd58de9095d1c0b88

Observation c69b76a0-2013-4a16-bcb6-6290905136b1 · outbound

This paper cites Unanswerable.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unanswerable

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.886834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:54:56.474300Z digest=sha256:9d995844ec8309354a1964df8c93596ce25cb20acdf648682fcfccf114dd8973

Pith citing papers

No inbound Pith citation observations are available.