Pith. sign in

Paper Citation Record · LEDGER

Performance Analysis of Traditional VQA Models Under Limited Computational Resources

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2502.05738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05738 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:38.360640Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93978d07-bc26-455b-8dd6-472c274992bb · outbound

This paper cites Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.619877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.244971Z digest=sha256:99af4795fae974fd3a20d0ef6b9162d9879c6e4afb6e289dd7a8ea60564334a2

Observation 54be280f-3706-4b34-9961-c3a2669e43b8 · outbound

This paper cites Learning Convolutional Text Representations for Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Learning Convolutional Text Representations for Visual Question Answering

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.594649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.260605Z digest=sha256:6e28ea8a95441e9e4a74a392893c3b686fe195bbb064f984ab60f621c0cf5d82

Observation 4ef1d1db-b1c5-4d74-9e93-0b166b93b7e1 · outbound

This paper cites Structured Attentions for Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Structured Attentions for Visual Question Answering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.579059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.265022Z digest=sha256:80f3b3b5bf30b0c48ab25826d576ad8a475c6e6e0daa113ee1d4334132fb9b38

Observation 36ec7fbb-05a3-481e-8938-e2da7d1e3f1c · outbound

This paper cites iVQA: Inverse Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources iVQA: Inverse Visual Question Answering

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.563465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.269298Z digest=sha256:c91eec24499c9bc6069044007bba5b702fc118c2ebca73319846f64286b7d6c8

Observation 0cbe02dc-f50d-47f4-b404-8ea3b1e126d1 · outbound

This paper cites Self-critical Sequence Training for Image Captioning.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Self-critical Sequence Training for Image Captioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.281868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.281868Z digest=sha256:1e26086715a424d84a4fe378a1efc62199275612905bd7862eb3e2e654203d9f

Observation d9b0ba65-3614-475f-a3e6-74181f19a943 · outbound

This paper cites X-Linear Attention Networks for Image Captioning.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources X-Linear Attention Networks for Image Captioning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.520079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.295039Z digest=sha256:8894b776aedd1d117d4a93ea137fca4f888d97d365e682709be8e17a63192e5a

Observation dee30146-85d3-47fd-849a-dbc2a2bdfa95 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.299321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.299321Z digest=sha256:8a2a1b01d1121d3352349dde84c26b0c7456ccb913ba2537aa0abab5f6b75aa5

Observation 99ef28dd-9510-4c9b-9edb-cf56bd770c80 · outbound

This paper cites Beyond a pre-trained object detector: Cross-modal textual and visual context for image captioning,.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Beyond a pre-trained object detector: Cross-modal textual and visual context for image captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:10:38.632483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.303290Z digest=sha256:a4b036752b8de166bef7d0a634252cb5f8c8def2795b1379737de7fae5ea2d0f

Observation 11fe739e-246c-4a01-afcd-f6a528289fe1 · outbound

This paper cites Transformer-Based Multi-modal Proposal and Re-Rank for Wikipedia Image-Caption Matching.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Transformer-Based Multi-modal Proposal and Re-Rank for Wikipedia Image-Caption Matching

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.494657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:10:38.306986Z digest=sha256:41fc814ffae6aeb2bfe03e80af5bfb19005e1e1f8bb45651143bce382efa2d90

Observation f6c2cb35-f96e-4aea-9037-75b6b9e78df8 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.311056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.311056Z digest=sha256:3208f931f7c59c5990c1ed96188ef645d8898c53b569e29d5f2785cf2a859a69

Observation 0f24e703-7482-4897-8b24-d8eca8e55813 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.315617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.315617Z digest=sha256:69bdc047c063cbaf54542bd8b109cd2375f4c600e3215b8565ff9549deeb4e1d

Observation ed3f86ac-51f9-49e3-a784-ef6b44ceaaf6 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Training data-efficient image transformers & distillation through attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.319474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.319474Z digest=sha256:b46eaa0b5ffca63de28d978ce5167e478216166d088b92d30b91e65e24668d44

Observation 5b875199-5513-4653-99ad-c25625dfa6a5 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.323443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.323443Z digest=sha256:28dce48a3459fc378b80f8452a9dd3d72181f14afb4485f4ec7464c72345585f

Observation e23c349d-d4e3-4c1b-b770-a7998624d357 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Distilling the Knowledge in a Neural Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.327827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.327827Z digest=sha256:ec7879cca76092ee989dccdc29dd5e3e24648f019c02482328b760ffb679227b

Observation d5d66f28-4c3d-4039-9f6d-cc525750cd45 · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources TinyBERT: Distilling BERT for Natural Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.356476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.356476Z digest=sha256:69fcf8a622bd2cdb00eff2d4a01881fcabab2d1267a6f754eb04dc254160ee24

Observation 3851e900-dd65-4cdd-8426-fa4fe8fdbd91 · outbound

This paper cites Learning both Weights and Connections for Efficient Neural Networks.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Learning both Weights and Connections for Efficient Neural Networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.335940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.335940Z digest=sha256:5b91a053734d06682599c5cd9936054f971e2188a53e533d3e7543ee8ad08865

Observation 50dcd689-b2e8-488a-af13-aaa88a2cbecd · outbound

This paper cites MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.353136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.353136Z digest=sha256:5747727c843cd424a9c02c4d0789fa775234c7108d87121cd9ba2b00b96e4f81

Observation d01917d4-6eed-4a4c-b8c1-de86dd67f13e · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.360640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.360640Z digest=sha256:cb548a7a8e2fee8a29eb4ddb00f751685bf48bd2f290b77cd3c38eced2b1d3e0

Observation f4d9f3a9-d543-4878-8903-1c708158c3f5 · outbound

This paper cites Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.278124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.278124Z digest=sha256:3b237adea25c6a1e1513e9083953407eeca8685f656328545c6629d8a49e2213

Observation 9d0941e8-7218-4156-b52d-027e66c44e7a · outbound

This paper cites Bidirectional Attention Flow for Machine Comprehension.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Bidirectional Attention Flow for Machine Comprehension

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.256348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.256348Z digest=sha256:5498a98c22e12d577090492db5f42543037bd318d17525133ce3927330f96636

Observation aa2525b9-c4fe-4f2a-90b8-d18bc7ab16ff · outbound

This paper cites Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.344144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.344144Z digest=sha256:e9367e724448e4fadb63b68f347eda2b334021d31dd798da0766cc3de42015a1

Observation 2385e11e-2888-4327-9185-b82d5e498178 · outbound

This paper cites Meshed-Memory Transformer for Image Captioning.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Meshed-Memory Transformer for Image Captioning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.290999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.290999Z digest=sha256:89c4e020c11ed2e9d5bc522fd6efdf8593891422217391b30fa2f9bcea5a8220

Pith citing papers

No inbound Pith citation observations are available.