Pith. sign in

Paper Citation Record · LEDGER

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

As of 21 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2608.07861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07861 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:19.424861Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

98 of 98 outbound references displayed

  • verified exact6
  • verified fuzzy62
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71d47ad8-bae6-495f-9981-555a93e09997 · outbound

This paper cites VQA: Visual Question Answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VQA: Visual Question Answering,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.017635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.017635Z digest=sha256:0061fceb1d4754809ac6abe97cf39a9ab80fa889a0b80dd96b6720453052f97c

Observation 8d7bdee0-3f56-4293-b5f7-efcfc3389f5c · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.022230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.022230Z digest=sha256:7aa2c904818acc051428fcf53732077d53355cdf2fd869ca03557a785e83f7eb

Observation aeee4825-6bf1-400e-b230-b03f9f69aee0 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.026832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.026832Z digest=sha256:1bfb350e933269951e14597c7830f58b959e31c5b7340c2f4ba53d0e3a2ac353

Observation 51d95b2f-c7d9-4ea4-819e-9c7989f2f0c0 · outbound

This paper cites Ray-Ban Meta AI glasses gen 2 & gen 1.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Ray-Ban Meta AI glasses gen 2 & gen 1

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.031728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.031728Z digest=sha256:0b14e2d72a985662baaa522014f701be44cd9f05a0b1f76bfc3cfc5900ae28db

Observation ca9a5527-1242-4b86-bd1c-85c3fadb4b4f · outbound

This paper cites Vision-language models for vision tasks: A survey,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Vision-language models for vision tasks: A survey,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.036101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.036101Z digest=sha256:935e23da71266c0747dbda169c4f13a6f6ea415e64252974f3e4e10df820345a

Observation c423ef81-b965-4ded-9423-26bd29476cc2 · outbound

This paper cites A survey on multimodal large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems A survey on multimodal large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.040708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.040708Z digest=sha256:e782e37086e779a7d83348f3649354cbcffd2e067f6e6fe43ad0f1971306deb3

Observation 5f6c1989-ff9c-411f-9b78-d5a98caf6067 · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VizWiz grand challenge: Answering visual questions from blind people,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.045319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.045319Z digest=sha256:6cf8a0f50f03ee0722ff912a7ab525f5a0ca9e2ea9e350b21238067851a9e7c0

Observation 2ec2a70d-2d04-40e7-9f53-bd449a390a50 · outbound

This paper cites Be My Eyes: Lend your eyes to the blind.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Be My Eyes: Lend your eyes to the blind

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.049342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.049342Z digest=sha256:3f94d5a8eb4eff9b6d25b7ad07480eab84f4aa4507641218a041b4d484a7b763

Observation c7a0d083-22a2-4034-b581-b6528d7cdf17 · outbound

This paper cites Long-form answers to visual questions from blind and low vision people,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Long-form answers to visual questions from blind and low vision people,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.053382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.053382Z digest=sha256:15b9955d857d2294e321ec11631c630158117587171722a888a95a9e613c3bdd

Observation f809fab8-781e-4a98-a07b-14dbb91febbb · outbound

This paper cites Augmented reality anatomy visualization for surgery assistance with HoloLens,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Augmented reality anatomy visualization for surgery assistance with HoloLens,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.057368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.057368Z digest=sha256:e35246d8774a76b25d1f023a18b6bf3a1c90df5ce8fbeb053ff9ef573fdd2849

Observation f4e53888-1f74-4bcb-8700-525f7fbc0d40 · outbound

This paper cites LLMs Enable Context-Aware Augmented Reality in Surgical Navigation.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems LLMs Enable Context-Aware Augmented Reality in Surgical Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.847198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.061282Z digest=sha256:2ff4bc374b9a5f8570e934a4ec0526b94369ca0b89942b8023eecf37cd343ba5

Observation 1f4ba1a6-b0cd-47fd-affe-6011636153dd · outbound

This paper cites DriVQA: A gaze- based dataset for visual question answering in driving scenarios,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems DriVQA: A gaze- based dataset for visual question answering in driving scenarios,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.065974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.065974Z digest=sha256:de7144ccfcfdec8bc8e459d1a8f154bfbbd31e51ce424b85d275797b7dec6c87

Observation f23872a6-04d1-484e-93c5-fa11a7b911cf · outbound

This paper cites Mimicking human attention in driving scenarios for enhanced visual question answering: Insights from eye-tracking and the human attention filter,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mimicking human attention in driving scenarios for enhanced visual question answering: Insights from eye-tracking and the human attention filter,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.069805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.069805Z digest=sha256:e93700e786f8c18d714dabfb267ec22937e3b03f965f8605da6de48f2ede14ec

Observation 0e174cce-37bf-44c7-a1d0-9ecab01484fb · outbound

This paper cites Demystifying small language models for edge deployment,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Demystifying small language models for edge deployment,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.073724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.073724Z digest=sha256:5d6c38db9c21b071fc68a2d58688efe81397108e047552e28e0c2e9ea6a38d07

Observation 4f3be074-5d14-4261-9b0f-284edc6d3958 · outbound

This paper cites Efficient processing of deep neural networks: A tutorial and survey,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Efficient processing of deep neural networks: A tutorial and survey,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.077650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.077650Z digest=sha256:248cddfa442be239caeda2fa5daa2c2b7ad9418e321ac680acee4f9e5fd5e764

Observation 3f716f81-e7ef-4153-bce3-1a2d210cacd2 · outbound

This paper cites LLM in a flash: Efficient large language model inference with limited memory,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems LLM in a flash: Efficient large language model inference with limited memory,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.082137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.082137Z digest=sha256:c7cb7149ce3f4d9c9bfcf3adb2440262c086a1c6ec3af03c25cd987c8daf74fc

Observation 093f752c-10f3-48fc-aea0-9e19b5827545 · outbound

This paper cites Mo- bileLLM: Optimizing sub-billion parameter language models for on- device use cases,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mo- bileLLM: Optimizing sub-billion parameter language models for on- device use cases,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.086239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.086239Z digest=sha256:b1d0fbd5ff4cef34ac2cc721efac5b3a09ec058ce3a9920c3056a8c51c50d369

Observation ed140e8c-6bb7-4a98-bae5-f052bb71ac1e · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.090469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.090469Z digest=sha256:1ae1a0231af7c712a392fb9a683f911c4697427d308c0752ac4f4fa4d9bbde0e

Observation 5591e1b1-915d-433f-9465-befe03319a4a · outbound

This paper cites Bench- marking tinyml systems: Challenges and direction,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Bench- marking tinyml systems: Challenges and direction,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.095087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.095087Z digest=sha256:efc3aa5d8932d98718466954a7c6067ae867e2d43f2fbfaaf93522ef29b06428

Observation e0624723-5114-4f44-ba70-eaa0a743094d · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.099492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.099492Z digest=sha256:c87496fb3bf2a0fbb2ab892bf24c5d3f9e7f4cd3cb0a0fc8ba34001eeb1eb023

Observation 030b6d04-1204-49be-9816-dcd0f2bb3b2d · outbound

This paper cites Rokid AI & AR glasses—redefining reality.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Rokid AI & AR glasses—redefining reality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.103858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.103858Z digest=sha256:3484c307464cf47c2fad2a37af397afa75177a70a0cd2e671750435fe4338832

Observation 33c982d4-3c83-47e6-acb7-fa530e2c4d11 · outbound

This paper cites Project Astra: A research prototype exploring the future of AI assistants.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Project Astra: A research prototype exploring the future of AI assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.108144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.108144Z digest=sha256:5bf5f1c3ff954d8f8790d4e880cf3ebdd6e2a810e5aa4a00abd5599ece4cb3bf

Observation 43286f77-d8ae-43c3-9bf8-3483d41a9953 · outbound

This paper cites Google vs meta smart glasses: Which AI frames are better.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Google vs meta smart glasses: Which AI frames are better

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.719204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.111993Z digest=sha256:fb44077d64332b1007365bc238e1b009e12c7f0567962c67a7e19714581c1398

Observation f4719216-9ab0-4982-9300-c7e5fadb17c7 · outbound

This paper cites Edge cloud offloading algorithms: Issues, methods, and perspectives,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Edge cloud offloading algorithms: Issues, methods, and perspectives,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.705801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.115924Z digest=sha256:c86ac033f21851013cb1f5ab968812f9f45eefa427913b1089de81169f255932

Observation e2197941-1bf2-423d-afa5-c6c52a9dda66 · outbound

This paper cites Cosmos:computation offloading as a service for mobile devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Cosmos:computation offloading as a service for mobile devices,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.690590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.119812Z digest=sha256:331035f7a7b35af03017024f7af36bba3d387b0b89905df34964c7107657f1d3

Observation b49ec1be-5003-4b1d-8980-a2abd19fa10c · outbound

This paper cites To offload or not to offload? the bandwidth and energy costs of mobile cloud comput- ing,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems To offload or not to offload? the bandwidth and energy costs of mobile cloud comput- ing,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.675452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.123655Z digest=sha256:43dff4ae439ffafb5a0d6aaf8812d3e0f2380244d988772e84de28c689702858

Observation 46a2ff98-5f5e-4710-8ec2-ac5fa2d36e6f · outbound

This paper cites Images and vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Images and vision,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.662673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.127468Z digest=sha256:8428fd4cdda1ae356a320bfeb390cac8d29719288f9c3e6b7ee7362597293f12

Observation e1199267-7777-484a-be17-82f4ba020a32 · outbound

This paper cites Understand and count tokens — gemini api,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Understand and count tokens — gemini api,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.650544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.131426Z digest=sha256:2b61fd47025327f4248cd4a89530a7e4838bb4ce0a949f222573bbe524c2fa3f

Observation e2f21d7e-8e38-48c8-a39c-c1f0178b018b · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems A-vit: Adaptive tokens for efficient vision transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.636667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.135656Z digest=sha256:ab2b2192cb3c457106d66372a9b58ea7978f2fcf9af0b6fd9732405d3accac45

Observation 6d2e4ea8-265e-463d-9465-ec7bf0e3941a · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganiza- tions,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Not all patches are what you need: Expediting vision transformers via token reorganiza- tions,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.139426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.139426Z digest=sha256:1e2aad577d9117c52b0e0fccde6079e33370d41f14a61fc7204a23d348b18c98

Observation b1f216e8-d0ab-496d-9366-3b51981bba7a · outbound

This paper cites Token merging: Your vit but faster,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Token merging: Your vit but faster,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.615091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.143097Z digest=sha256:97735a5d3ead90f8a0a2337528a901bd6ec388aeda1f445cd8b4dc414d169086

Observation 6fb6ed1d-b647-4032-bb55-085660552667 · outbound

This paper cites Elf: Accelerate high-resolution mobile deep vision with content-aware parallel offloading,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Elf: Accelerate high-resolution mobile deep vision with content-aware parallel offloading,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.601893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.146721Z digest=sha256:f72048bccd5fc9602483efb5d902732c4fd0c119f6a47fccf93832245e99364d

Observation c815e248-ed3e-4b28-ba44-419e630d7ae0 · outbound

This paper cites Accumo: Accuracy- centric multitask offloading in edge-assisted mobile augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Accumo: Accuracy- centric multitask offloading in edge-assisted mobile augmented reality,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.589138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.150710Z digest=sha256:24a23d096948c6ca95fc54794cff010e1eae4adf24b0ca7b9826c89a28534ef7

Observation db119f59-071d-4edf-b611-a1ad706dad96 · outbound

This paper cites Deep contextualized compressive offloading for images,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep contextualized compressive offloading for images,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.575387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.154686Z digest=sha256:e55ed3e14c553d66959f4c6a932391097d3acde7be079d35486e8f663221e146

Observation 62d30240-4617-49fe-8b99-a546eb9e352a · outbound

This paper cites Towards wearable cognitive assistance,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Towards wearable cognitive assistance,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.561858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.158591Z digest=sha256:dd0da8a289797849277afd1aa8c0385b8358f5cdabdb68bf93837628a937f48f

Observation e1736367-04fd-4dca-81ad-2fa144d30a94 · outbound

This paper cites Deep learning with edge computing: A review,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep learning with edge computing: A review,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.162754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.162754Z digest=sha256:3b62e782234d450f7077baf3eb0e9cde02ed466d620e3ab5357d82cf18b1a301

Observation b611c525-e955-40b7-9ec3-84c34dd7569d · outbound

This paper cites Color-to-grayscale: Does the method matter in image recognition?,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Color-to-grayscale: Does the method matter in image recognition?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.540375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.167041Z digest=sha256:fe4352f1da52cf4a4c9d0e92e97b3787dd92b0251fb3f7e2f325652017a2159a

Observation 82072b31-de42-4ac8-bb17-d7543c2d9d70 · outbound

This paper cites The jpeg still picture compression standard,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems The jpeg still picture compression standard,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.170915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.170915Z digest=sha256:9e1462fe1465ac594b19a0d9366e8b4b6e4468cae178dfbd5a63213a2ba656fa

Observation a4b2b770-2d6d-4fd1-af7e-2750e9d6511e · outbound

This paper cites Learning to resize images for computer vision tasks,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Learning to resize images for computer vision tasks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.520409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.174922Z digest=sha256:109e5c99b84e7cc974c7a816cf5147c7a828e3f2484f8ba857e3043c690e3288

Observation 6e7aa746-6bc4-4004-a37f-d48474ab831a · outbound

This paper cites VQA-MHUG: A gaze dataset to study multimodal neural attention in visual question answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VQA-MHUG: A gaze dataset to study multimodal neural attention in visual question answering,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.509216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.178967Z digest=sha256:56009da4e48180ea7f4e023e8e8d686ac3580f0d34c2f710a5483ed741c7a162

Observation 8c632c78-07be-4139-a733-c27e553b0ce6 · outbound

This paper cites Eye gaze tells you where to compute: Gaze-driven efficient vlms,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Eye gaze tells you where to compute: Gaze-driven efficient vlms,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.497438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.182640Z digest=sha256:f582cdbad0c795b9a4ec6d9173913b132af8cc9c79cc7422e89f4b28e8eacb00

Observation ea0c5c9b-871a-45d3-924c-b67571e48b83 · outbound

This paper cites Saliency detection: A spectral residual approach,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency detection: A spectral residual approach,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.485099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.186564Z digest=sha256:8a46cde98e178b7a7eb766550aed4d6c7c9c8b6120dc9825c49022f08d0c6a6b

Observation 2984d488-75ea-4a56-b6c8-3cfd71b74027 · outbound

This paper cites Saliency driven perceptual image compression,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency driven perceptual image compression,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.473103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.190484Z digest=sha256:0985fa64ecca229a003d84c30d9f7db8f56b73cb233b3087e6c6c2bd3884b022

Observation 46657777-0be2-4c24-a490-16df0595a874 · outbound

This paper cites Visual cropping improves zero-shot question answering of multimodal large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Visual cropping improves zero-shot question answering of multimodal large language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.461376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.195307Z digest=sha256:349f8a3b9795ed12f09ee9729e1a1d42b530fb07b2f164f802319e5270c116e6

Observation 0661ec84-8aed-4a0c-9051-4f3a6796db86 · outbound

This paper cites Edge assisted real-time object detection for mobile augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Edge assisted real-time object detection for mobile augmented reality,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.450028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.199136Z digest=sha256:2e75fd04cc1f649510dfb531c697dce49f87a99b595764236d04241898eef321

Observation 6d1ff298-076a-42f4-9024-d3724fdc2cd0 · outbound

This paper cites Glimpse: Continuous, real-time object recognition on mobile devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Glimpse: Continuous, real-time object recognition on mobile devices,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.438645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.202934Z digest=sha256:179caf1c4c96839608ce7b621520c7b8d2c33e628c158d2975051dd389fdf20e

Observation bb99a795-38e0-4c1d-834e-d84387269ab0 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.426348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.207129Z digest=sha256:3888a969f9530487a9b78b59f1958befbe271fe96c806a16c37e4e7a8aac49e3

Observation a4b4c0c5-25ef-4e90-8c8c-a85746a2e407 · outbound

This paper cites MMBench: Is your multi- modal model an all-around player?,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MMBench: Is your multi- modal model an all-around player?,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.414074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.211247Z digest=sha256:0c3dba2a70bc46b101ff2b51b9e259706d0411b8dd57b6efbbfa8dea20787e91

Observation af7e46f8-ecda-49b8-af04-3f441fd59c3e · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.401510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.215889Z digest=sha256:99496d54dcf44c3edc4c4d843a3524ca6512e433a4153e40fed46438bd6e137b

Observation 2659ab9f-0bc5-4bfb-bbd5-c13dde02c5d8 · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.389142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.219803Z digest=sha256:0daaa8bf857577cdd8192c3a7b5504ef27f810f94bcb01782d899dc63282a493

Observation fab0333c-8dcf-4f0b-a835-2e2417777199 · outbound

This paper cites HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.376588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.223759Z digest=sha256:e9adce563a464c5f1c78db0c21f2cee057294e0d186fbd51d6e2ac7729aea027

Observation 92215827-6b85-4a43-9454-1c9eaaba8522 · outbound

This paper cites WearVQA: A visual question answering benchmark for wearables in egocentric authentic real-world scenarios,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems WearVQA: A visual question answering benchmark for wearables in egocentric authentic real-world scenarios,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.364074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.228109Z digest=sha256:125d5d9b5983c56337d3e3153925bbc423acdaee85d49b3264f99df32d193aba

Observation ea3fede4-c24d-4cf7-a57f-7641895da963 · outbound

This paper cites SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.232005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.232005Z digest=sha256:cf6a3f0ec2c09650b372b63dc3f361ecff3295ff8871fc618bd975688366731d

Observation 9510abb9-44a5-4dd3-aab4-f6334dcd9568 · outbound

This paper cites CRAG-MM: Multi-modal multi-turn comprehen- sive RAG benchmark,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems CRAG-MM: Multi-modal multi-turn comprehen- sive RAG benchmark,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.238653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.238653Z digest=sha256:d1dcacb1ae4c40f91d700bec35d261c000367a103d9d624216a395aa1c202600

Observation d117e68e-2597-4511-8a8f-cd4b438c1b90 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems OK-VQA: A visual question answering benchmark requiring external knowledge,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.351783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.242494Z digest=sha256:32cf4783049f684c09af7daec2ee12532d9c6875251243c207bf9b0ba51e5f81

Observation a41e1861-e52a-4ab1-ba2f-2da8df2db08a · outbound

This paper cites MISAR: A multimodal instructional system with augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MISAR: A multimodal instructional system with augmented reality,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.339738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.246484Z digest=sha256:b2cc77c8215c23bebcf7a3a2056d9ecf9d7fa90291acddf077207ada370eb1b6

Observation e10d0e33-8929-4845-b0c5-9b401e07c111 · outbound

This paper cites Guided Reality: Generating visually-enriched AR task guidance with LLMs and vision models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Guided Reality: Generating visually-enriched AR task guidance with LLMs and vision models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.328278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.250850Z digest=sha256:05ca1e2d3141b8bdcccc2c48f6f840062fa4988e1a2aebdf2aa6a2503c22d2a9

Observation 54592885-ffbf-410c-9493-1fea13cc1cd7 · outbound

This paper cites EmBARDiment: an Embodied AI Agent for Productivity in XR.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EmBARDiment: an Embodied AI Agent for Productivity in XR

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.254659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.254659Z digest=sha256:7124b4b0df59d5247b3cebf33f492535ce7d862b49e643da9ad6cc6ee90ef401

Observation ef168b0b-05ee-4c8c-acfc-835114242bdf · outbound

This paper cites GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.315950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.261066Z digest=sha256:481c04402998305016019d43ba9696111ab28ec0a7bdd5f6a218d77144a45f96

Observation 743d7ba5-8ffe-416b-9a38-98f7ed88fb16 · outbound

This paper cites Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.701778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.265562Z digest=sha256:3ffcb4216e0d5e25c0cf5e6d1a290ce873c2aa16d362533b65437f44b1a81a1a

Observation 01dd9cdc-092c-4bb0-bc95-48c813e8338f · outbound

This paper cites Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.676748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.270631Z digest=sha256:5e1f1f76cc950bc67b5fb5b3c8bfa910bb1202ebc21152da2af8ee87ec4b51cc

Observation 6676f53c-2d73-4ea9-8c37-0990526deebc · outbound

This paper cites Exploring the use of VLMs for navigation assistance for people with blindness and low vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Exploring the use of VLMs for navigation assistance for people with blindness and low vision,

Reference 62

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:50:19.657100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.275112Z digest=sha256:627c4e32a609b259aceace4158430364170b7a33fbf12f863a87b8091173fcb6

Observation ed1391d9-df16-421e-b77a-79a1b2965b96 · outbound

This paper cites BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.569534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.280108Z digest=sha256:ca9bf75d6964d9eea8581c2109a6975dd2636ebef18b3c7c16e477ca3215d41b

Observation 22753e43-642f-43cd-833a-286ca03f7fa5 · outbound

This paper cites Objective assessment of the webp image coding algorithm,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Objective assessment of the webp image coding algorithm,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.304358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.284581Z digest=sha256:0cb2279ff89388e2a08ca1035f57d97ed1e67a14a5056e3f65ff80fb16270050

Observation 01bace46-f84d-4361-85c5-794780a21227 · outbound

This paper cites Vision - claude api docs,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Vision - claude api docs,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.292984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.288751Z digest=sha256:38d654eb261fac2b14d1132fdd00f23e8a5c617520eb783bc7de6f70bfa63b98

Observation fa6fb7df-483a-4d21-956c-805c2a386197 · outbound

This paper cites Image understanding — gemini api,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Image understanding — gemini api,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.281030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.292821Z digest=sha256:b38041f08f875d7ded235a1b708125a08cf09dcf860c7056cd368ca13458b015

Observation 41af0a65-9a0f-4dd8-8072-cab64b6cb2a6 · outbound

This paper cites Image interpolation and resam- pling,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Image interpolation and resam- pling,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.268208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.297025Z digest=sha256:84fd65280bf35f0484fb31a5596dde05e2308c6fe1be846216d5c67ff091723b

Observation f99b59f9-f810-4d14-8e34-e9a36a3fd0cd · outbound

This paper cites Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.256188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.301374Z digest=sha256:8d1a4c757c69d6c9f081250fa386209e94ca820ec8e78603d5ea7ad217749d0b

Observation 6c8b19a3-96f7-42bb-86d6-9061eba41935 · outbound

This paper cites V oila-a: Aligning vision-language models with user's gaze attention,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems V oila-a: Aligning vision-language models with user's gaze attention,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.242817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.306560Z digest=sha256:527c052a7e4fa3141496310df5c99bc28d9590ef254889738dd8a94ca13b1701

Observation f87e2bc5-b3fd-4be5-8c95-6ba2989c591d · outbound

This paper cites EgoVQA—an egocentric video question answering benchmark dataset,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EgoVQA—an egocentric video question answering benchmark dataset,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.228121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.310732Z digest=sha256:a728aa2cd308c32969d601b98c3384f689d9df791452314722fccf87ec198e40

Observation 9bf1c15d-5ff4-405f-8e17-d64dfc0701e2 · outbound

This paper cites Neural machine translation of rare words with subword units,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Neural machine translation of rare words with subword units,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.216201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.314944Z digest=sha256:45a777d09005137545e0d3ce5b2829b1354f8c771cd33d4c9b743cba338563b4

Observation 2fd64ec6-85d1-4967-b5d8-38c431fa085d · outbound

This paper cites What are tokens and how to count them?.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems What are tokens and how to count them?

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.204919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.319134Z digest=sha256:193a2ee10572cd641078734b4add0d283494b208b59d0f3cc4be1a0b8a082a98

Observation 5da132dd-8582-4cbc-8c93-f29ca696bc7f · outbound

This paper cites Saliency based image crop- ping,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency based image crop- ping,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.192706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.323268Z digest=sha256:fb8b41bffd1efa640c5956d2389c22717070275de6573ae62dd83fb0102e9e94

Observation 2794bdd8-29c7-47b4-a08d-47226a09f943 · outbound

This paper cites Benchmarking deep learning models for object detection on edge computing devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Benchmarking deep learning models for object detection on edge computing devices,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.179180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.327140Z digest=sha256:6f97582417b0e8c3b179f0a5a4d5f5333c7bb2ddb45ed8b696e531649742c213

Observation 89aca3b1-210a-4f3c-a678-ffd30360e0e6 · outbound

This paper cites Region-of-interest extraction method to increase object-detection performance in remote monitoring system,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Region-of-interest extraction method to increase object-detection performance in remote monitoring system,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.166519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.330909Z digest=sha256:5e37955908ff3fa269046bdbd25f3e115319e56ed03abf7a87695a9885f3d386

Observation d50df610-97c5-461a-ad07-33e8fbcfb191 · outbound

This paper cites Improving automatic VQA evaluation using large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Improving automatic VQA evaluation using large language models,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.154479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.335462Z digest=sha256:dcd933a6f37f77b28f585acc80772504fc5c4f0ffd53d65f1152457d4f65965c

Observation 33234788-e51a-4184-b466-9825241e5b41 · outbound

This paper cites Mind the uncertainty in human disagreement: Evaluating discrepancies between model predictions and human responses in vqa,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mind the uncertainty in human disagreement: Evaluating discrepancies between model predictions and human responses in vqa,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.142051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.339617Z digest=sha256:5013f4768caf57e7684a23ff54e037d7948c0f9f17a9813d3451c73c0e62c41a

Observation 0f4ff4da-137d-44ba-8cf0-faab0953198f · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.128003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.344182Z digest=sha256:db736eeb661c2237c7f378d7b034b441ce2c10a774cd02598a4b13fec5c7c238

Observation afd1f3e0-03c3-4560-98cf-b74131d5675a · outbound

This paper cites G-Eval: NLG evaluation using GPT-4 with better human alignment,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems G-Eval: NLG evaluation using GPT-4 with better human alignment,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.113194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.348271Z digest=sha256:04ed9d0cbe33eb633ee96839bb6ea1e49a0a33caf4b094bfdc1cfffb350199c3

Observation 57d015c5-2614-4081-8fbc-ef1c06409793 · outbound

This paper cites Chatar: Conversation support using large language model and augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Chatar: Conversation support using large language model and augmented reality,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.099459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.352344Z digest=sha256:718b3254ab510a98b58578c8a2af5e02a1afb13af51057a2221a09a5f3877577

Observation 0db091eb-cd22-4e25-9b2a-5515d36fb4b3 · outbound

This paper cites Next-generation networking and edge computing for mixed reality real-time interactive systems,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Next-generation networking and edge computing for mixed reality real-time interactive systems,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.086379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.356200Z digest=sha256:4fc3522e9f7d3220b3c89555a6d4f13bb3965a62c023f126871b587fe40b09d3

Observation 43a2847e-f2a4-4cdc-866c-3ed4e713706d · outbound

This paper cites An egocentric vision-language model based portable real-time smart assistant,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems An egocentric vision-language model based portable real-time smart assistant,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.071213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.360215Z digest=sha256:3290569f6db1ec750b42d63c180cb8c27cc6b6a1355fcd5163635e682ebefc30

Observation 4cbd87af-581d-45c3-a2de-fb96d7c5dbfa · outbound

This paper cites Teaching LLMs to see and guide: Context-aware real-time assistance in augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Teaching LLMs to see and guide: Context-aware real-time assistance in augmented reality,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.364187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.364187Z digest=sha256:3353186097faea0b97a31dd2bfd9a0b5355fc7680a5883b355b5305faf25fd05

Observation 3999f9fc-18db-49bc-85a5-767882c1d7fb · outbound

This paper cites Gartner predicts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gartner predicts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.057437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.368841Z digest=sha256:b6f41ea34a4fe1ffec9b88482d94b49bfdf16efb8b3d82138526c58d3a4b3556

Observation 9ad2d767-1a29-49bb-8fc9-b5b4f177566a · outbound

This paper cites Deep learning-based object detection in augmented reality: A systematic review,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep learning-based object detection in augmented reality: A systematic review,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.372822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.372822Z digest=sha256:a2891a17ccd5fdd6a8921c7354ee631a4614e2d52c7f5c0eedd884b36c7ecb81

Observation b3c695f9-5789-4416-885c-8d6bf76902b8 · outbound

This paper cites Realtime API with WebRTC,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Realtime API with WebRTC,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.036424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.376497Z digest=sha256:e2a7bb8eef64f0e25d27ed39954eec4058deb90d28c190cc512b403a595e217a

Observation 2437408b-3e1a-4cee-9a4e-18cefaaabc6c · outbound

This paper cites Realtime API with WebSocket,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Realtime API with WebSocket,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.022217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.380300Z digest=sha256:f5531c552d345f9c269dbbe37cd32208c0837bbf006cfa1a303257a51409e857

Observation 0b4103ae-d36f-4f62-b6d0-c6c63b65fa68 · outbound

This paper cites Bench- marking the effects of operating system interference on extreme-scale parallel machines,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Bench- marking the effects of operating system interference on extreme-scale parallel machines,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.006413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.384617Z digest=sha256:d21f58ecd565012204e5414b81823c44240b90accb4c997ac3e6103c3b5458dd

Observation 1f047041-fa3f-4727-9554-0090e9e3532e · outbound

This paper cites Understanding the causes of performance variability in HPC workloads,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Understanding the causes of performance variability in HPC workloads,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.990984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.389165Z digest=sha256:563e02a44eadb4414212b2bc2ae293c65028b69ed65be705e776bfa189e4c353

Observation 129c06b6-3e73-46a7-b6b0-c9eac695ac01 · outbound

This paper cites Overhead Measurement Noise in Different Runtime Environments.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Overhead Measurement Noise in Different Runtime Environments

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.466260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.392944Z digest=sha256:c4272551b813571dd257a6cef03d4987f02bc6031308b0995f2874c80442e2d3

Observation b5e9adb5-74f6-44ab-8670-c8f301c10ab1 · outbound

This paper cites Using microbenchmarks to evaluate system performance,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Using microbenchmarks to evaluate system performance,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.976627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.397014Z digest=sha256:a6572ebafb3a5bd8a5bda41942cbb127afbfb987aeae1fcb86eb2f4faad83ab4

Observation 308695cf-3a57-4dd8-945a-a254dd5502f4 · outbound

This paper cites Beyond inference: Performance analysis of DNN server overheads for computer vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Beyond inference: Performance analysis of DNN server overheads for computer vision,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.962037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.400903Z digest=sha256:7b43fdc042057bc61af786804fcdc8777e60724c3c69096d580fa35fc562d60d

Observation 199f6712-add6-4936-95dd-b41d3f3fd1a6 · outbound

This paper cites EdgeYOLO: An edge-real- time object detector,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EdgeYOLO: An edge-real- time object detector,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.948887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.405290Z digest=sha256:d7d507b08abb9febf8bb422e9a3ca2c1f961877e7525dd648adc1ba45a7443ad

Observation 464bb5f5-a174-40fb-860b-14cfe70ae457 · outbound

This paper cites Images and vision — calculating image tokens.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Images and vision — calculating image tokens

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.935016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.409358Z digest=sha256:80c6fc8d3548f04c0a32141a5cd0ad7ccc6f60e776b679254370c3043765acbd

Observation 7d4d740a-2e86-46f3-b44f-b283bba3fcf4 · outbound

This paper cites Per-part media resolution (Gemini 3 only).

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Per-part media resolution (Gemini 3 only)

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.920442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.413398Z digest=sha256:20b52da625fd08f81e354fb55f53e7fd033c25b23a2a1ab497ea3c7667a9eb9b

Observation 279c4101-1146-4c54-a220-213d08828304 · outbound

This paper cites Gemini 3 Flash — model documentation.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gemini 3 Flash — model documentation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.904265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.417282Z digest=sha256:1b9b1cd8a9e2e7bf13b2e56135d1aac710ed2b6d7a33ae66119b81f517ad18dd

Observation 6e963d5a-b3a1-4dab-97b1-0a1dc045532d · outbound

This paper cites Gemini 3 developer guide — generateContent API.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gemini 3 developer guide — generateContent API

Reference 97

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:50:19.877165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.424861Z digest=sha256:3dc4c47a2c3938cc315e84f7fd6770f29c73e054bdc78823df0eb86a82c9a950

Observation b0674577-89e0-4df8-88f0-db6abe44eb2d · outbound

This paper cites Accessed: 2026-05-28.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Accessed: 2026-05-28

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.891144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:50:19.421183Z digest=sha256:c078123a3a3663697a16cb386a616c13ecf50bbb7dc279903cb5efcc6ebcf7be

Pith citing papers

No inbound Pith citation observations are available.