Pith. sign in

Paper Citation Record · LEDGER

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

As of 21 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2608.07861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07861 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:19.424861Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

98 of 98 outbound references displayed

  • verified exact6
  • verified fuzzy62
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71d47ad8-bae6-495f-9981-555a93e09997 · outbound

This paper cites VQA: Visual Question Answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VQA: Visual Question Answering,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.017635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.017635Z digest=sha256:0061fceb1d4754809ac6abe97cf39a9ab80fa889a0b80dd96b6720453052f97c

Observation 8d7bdee0-3f56-4293-b5f7-efcfc3389f5c · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.022230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.022230Z digest=sha256:7aa2c904818acc051428fcf53732077d53355cdf2fd869ca03557a785e83f7eb

Observation aeee4825-6bf1-400e-b230-b03f9f69aee0 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.026832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.026832Z digest=sha256:1bfb350e933269951e14597c7830f58b959e31c5b7340c2f4ba53d0e3a2ac353

Observation 51d95b2f-c7d9-4ea4-819e-9c7989f2f0c0 · outbound

This paper cites Ray-Ban Meta AI glasses gen 2 & gen 1.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Ray-Ban Meta AI glasses gen 2 & gen 1

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.031728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.031728Z digest=sha256:0b14e2d72a985662baaa522014f701be44cd9f05a0b1f76bfc3cfc5900ae28db

Observation ca9a5527-1242-4b86-bd1c-85c3fadb4b4f · outbound

This paper cites Vision-language models for vision tasks: A survey,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Vision-language models for vision tasks: A survey,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.036101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.036101Z digest=sha256:935e23da71266c0747dbda169c4f13a6f6ea415e64252974f3e4e10df820345a

Observation c423ef81-b965-4ded-9423-26bd29476cc2 · outbound

This paper cites A survey on multimodal large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems A survey on multimodal large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.040708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.040708Z digest=sha256:e782e37086e779a7d83348f3649354cbcffd2e067f6e6fe43ad0f1971306deb3

Observation 5f6c1989-ff9c-411f-9b78-d5a98caf6067 · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VizWiz grand challenge: Answering visual questions from blind people,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.045319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.045319Z digest=sha256:6cf8a0f50f03ee0722ff912a7ab525f5a0ca9e2ea9e350b21238067851a9e7c0

Observation 2ec2a70d-2d04-40e7-9f53-bd449a390a50 · outbound

This paper cites Be My Eyes: Lend your eyes to the blind.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Be My Eyes: Lend your eyes to the blind

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.049342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.049342Z digest=sha256:3f94d5a8eb4eff9b6d25b7ad07480eab84f4aa4507641218a041b4d484a7b763

Observation c7a0d083-22a2-4034-b581-b6528d7cdf17 · outbound

This paper cites Long-form answers to visual questions from blind and low vision people,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Long-form answers to visual questions from blind and low vision people,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.053382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.053382Z digest=sha256:15b9955d857d2294e321ec11631c630158117587171722a888a95a9e613c3bdd

Observation f809fab8-781e-4a98-a07b-14dbb91febbb · outbound

This paper cites Augmented reality anatomy visualization for surgery assistance with HoloLens,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Augmented reality anatomy visualization for surgery assistance with HoloLens,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.057368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.057368Z digest=sha256:e35246d8774a76b25d1f023a18b6bf3a1c90df5ce8fbeb053ff9ef573fdd2849

Observation f4e53888-1f74-4bcb-8700-525f7fbc0d40 · outbound

This paper cites LLMs Enable Context-Aware Augmented Reality in Surgical Navigation.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems LLMs Enable Context-Aware Augmented Reality in Surgical Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.847198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.061282Z digest=sha256:2e692e2f223984fefb7af09a16c9224f53a562c632a3df572b96a9d8afd60f7f

Observation 1f4ba1a6-b0cd-47fd-affe-6011636153dd · outbound

This paper cites DriVQA: A gaze- based dataset for visual question answering in driving scenarios,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems DriVQA: A gaze- based dataset for visual question answering in driving scenarios,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.065974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.065974Z digest=sha256:de7144ccfcfdec8bc8e459d1a8f154bfbbd31e51ce424b85d275797b7dec6c87

Observation f23872a6-04d1-484e-93c5-fa11a7b911cf · outbound

This paper cites Mimicking human attention in driving scenarios for enhanced visual question answering: Insights from eye-tracking and the human attention filter,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mimicking human attention in driving scenarios for enhanced visual question answering: Insights from eye-tracking and the human attention filter,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.069805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.069805Z digest=sha256:e93700e786f8c18d714dabfb267ec22937e3b03f965f8605da6de48f2ede14ec

Observation 0e174cce-37bf-44c7-a1d0-9ecab01484fb · outbound

This paper cites Demystifying small language models for edge deployment,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Demystifying small language models for edge deployment,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.073724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.073724Z digest=sha256:5d6c38db9c21b071fc68a2d58688efe81397108e047552e28e0c2e9ea6a38d07

Observation 4f3be074-5d14-4261-9b0f-284edc6d3958 · outbound

This paper cites Efficient processing of deep neural networks: A tutorial and survey,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Efficient processing of deep neural networks: A tutorial and survey,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.077650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.077650Z digest=sha256:248cddfa442be239caeda2fa5daa2c2b7ad9418e321ac680acee4f9e5fd5e764

Observation 3f716f81-e7ef-4153-bce3-1a2d210cacd2 · outbound

This paper cites LLM in a flash: Efficient large language model inference with limited memory,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems LLM in a flash: Efficient large language model inference with limited memory,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.082137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.082137Z digest=sha256:c7cb7149ce3f4d9c9bfcf3adb2440262c086a1c6ec3af03c25cd987c8daf74fc

Observation 093f752c-10f3-48fc-aea0-9e19b5827545 · outbound

This paper cites Mo- bileLLM: Optimizing sub-billion parameter language models for on- device use cases,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mo- bileLLM: Optimizing sub-billion parameter language models for on- device use cases,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.086239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.086239Z digest=sha256:b1d0fbd5ff4cef34ac2cc721efac5b3a09ec058ce3a9920c3056a8c51c50d369

Observation ed140e8c-6bb7-4a98-bae5-f052bb71ac1e · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.090469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.090469Z digest=sha256:538a2c70acd4226196407d7469a1dd102d94a28b52c03f187e2d7aef9de491aa

Observation 5591e1b1-915d-433f-9465-befe03319a4a · outbound

This paper cites Bench- marking tinyml systems: Challenges and direction,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Bench- marking tinyml systems: Challenges and direction,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.095087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.095087Z digest=sha256:efc3aa5d8932d98718466954a7c6067ae867e2d43f2fbfaaf93522ef29b06428

Observation e0624723-5114-4f44-ba70-eaa0a743094d · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.099492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.099492Z digest=sha256:c87496fb3bf2a0fbb2ab892bf24c5d3f9e7f4cd3cb0a0fc8ba34001eeb1eb023

Observation 030b6d04-1204-49be-9816-dcd0f2bb3b2d · outbound

This paper cites Rokid AI & AR glasses—redefining reality.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Rokid AI & AR glasses—redefining reality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.103858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.103858Z digest=sha256:3484c307464cf47c2fad2a37af397afa75177a70a0cd2e671750435fe4338832

Observation 33c982d4-3c83-47e6-acb7-fa530e2c4d11 · outbound

This paper cites Project Astra: A research prototype exploring the future of AI assistants.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Project Astra: A research prototype exploring the future of AI assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.108144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.108144Z digest=sha256:5bf5f1c3ff954d8f8790d4e880cf3ebdd6e2a810e5aa4a00abd5599ece4cb3bf

Observation 43286f77-d8ae-43c3-9bf8-3483d41a9953 · outbound

This paper cites Google vs meta smart glasses: Which AI frames are better.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Google vs meta smart glasses: Which AI frames are better

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.719204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.111993Z digest=sha256:44a98632ce7fba127263f2bf4a0737ea1fdc5f12bb74a2efd464fe693fe1c32c

Observation f4719216-9ab0-4982-9300-c7e5fadb17c7 · outbound

This paper cites Edge cloud offloading algorithms: Issues, methods, and perspectives,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Edge cloud offloading algorithms: Issues, methods, and perspectives,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.705801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.115924Z digest=sha256:388aa574ebe7edd65bc8d519558fc7d01683720b4ba2c102f9341fcef9018959

Observation e2197941-1bf2-423d-afa5-c6c52a9dda66 · outbound

This paper cites Cosmos:computation offloading as a service for mobile devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Cosmos:computation offloading as a service for mobile devices,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.690590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.119812Z digest=sha256:5027308a54f5d0f971b579985f7fad836e5e13566b807ab0d9ba5513bf28466e

Observation b49ec1be-5003-4b1d-8980-a2abd19fa10c · outbound

This paper cites To offload or not to offload? the bandwidth and energy costs of mobile cloud comput- ing,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems To offload or not to offload? the bandwidth and energy costs of mobile cloud comput- ing,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.675452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.123655Z digest=sha256:b7cec591147d510157828a32eb15249a95481e6636c832e80dbde28f8a22731d

Observation 46a2ff98-5f5e-4710-8ec2-ac5fa2d36e6f · outbound

This paper cites Images and vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Images and vision,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.662673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.127468Z digest=sha256:71fa9e828631f409fc93ca2dd9711b8d02f506d69d0cc9c6996373546e0db817

Observation e1199267-7777-484a-be17-82f4ba020a32 · outbound

This paper cites Understand and count tokens — gemini api,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Understand and count tokens — gemini api,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.650544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.131426Z digest=sha256:3cd982054004b4215fd622f22de194849e3ff61e50ccd904c1439bfdb6b50470

Observation e2f21d7e-8e38-48c8-a39c-c1f0178b018b · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems A-vit: Adaptive tokens for efficient vision transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.636667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.135656Z digest=sha256:706a9b1af99bd9964f090796031383be68784fb0629c317f196de771ff3f404e

Observation 6d2e4ea8-265e-463d-9465-ec7bf0e3941a · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganiza- tions,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Not all patches are what you need: Expediting vision transformers via token reorganiza- tions,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.139426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.139426Z digest=sha256:1e2aad577d9117c52b0e0fccde6079e33370d41f14a61fc7204a23d348b18c98

Observation b1f216e8-d0ab-496d-9366-3b51981bba7a · outbound

This paper cites Token merging: Your vit but faster,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Token merging: Your vit but faster,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.615091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.143097Z digest=sha256:d28034433fe4b5d4ec98f177da467fb5d780350e26c32d3671421e915e48edfb

Observation 6fb6ed1d-b647-4032-bb55-085660552667 · outbound

This paper cites Elf: Accelerate high-resolution mobile deep vision with content-aware parallel offloading,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Elf: Accelerate high-resolution mobile deep vision with content-aware parallel offloading,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.601893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.146721Z digest=sha256:160bde3c3fd9510d13d71ae890110c5319b6d78b4d3cc8ee6e759aa26f6b4149

Observation c815e248-ed3e-4b28-ba44-419e630d7ae0 · outbound

This paper cites Accumo: Accuracy- centric multitask offloading in edge-assisted mobile augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Accumo: Accuracy- centric multitask offloading in edge-assisted mobile augmented reality,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.589138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.150710Z digest=sha256:be195aa57c09d9e0c7225a22f1d2347126949a3bf900886c2ae0f64a48c26e43

Observation db119f59-071d-4edf-b611-a1ad706dad96 · outbound

This paper cites Deep contextualized compressive offloading for images,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep contextualized compressive offloading for images,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.575387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.154686Z digest=sha256:41b8eea68a6e50b1f5f19695a2a61898ff7425d4eca4726f21fb4467be69ebc4

Observation 62d30240-4617-49fe-8b99-a546eb9e352a · outbound

This paper cites Towards wearable cognitive assistance,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Towards wearable cognitive assistance,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.561858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.158591Z digest=sha256:96ac6dd1fea725758b5aa181b9889afdf7355e50a3f735cf26856f6148eae6da

Observation e1736367-04fd-4dca-81ad-2fa144d30a94 · outbound

This paper cites Deep learning with edge computing: A review,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep learning with edge computing: A review,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.162754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.162754Z digest=sha256:3b62e782234d450f7077baf3eb0e9cde02ed466d620e3ab5357d82cf18b1a301

Observation b611c525-e955-40b7-9ec3-84c34dd7569d · outbound

This paper cites Color-to-grayscale: Does the method matter in image recognition?,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Color-to-grayscale: Does the method matter in image recognition?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.540375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.167041Z digest=sha256:b3dbc5852305ad4c66aa24741d96ca365b81803b8c87762d8f46155ea934508b

Observation 82072b31-de42-4ac8-bb17-d7543c2d9d70 · outbound

This paper cites The jpeg still picture compression standard,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems The jpeg still picture compression standard,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.170915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.170915Z digest=sha256:9e1462fe1465ac594b19a0d9366e8b4b6e4468cae178dfbd5a63213a2ba656fa

Observation a4b2b770-2d6d-4fd1-af7e-2750e9d6511e · outbound

This paper cites Learning to resize images for computer vision tasks,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Learning to resize images for computer vision tasks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.520409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.174922Z digest=sha256:2ec8c172e448662efb7b8adad1a0de83d7dd4cce5a79f5628b65d308d3896583

Observation 6e7aa746-6bc4-4004-a37f-d48474ab831a · outbound

This paper cites VQA-MHUG: A gaze dataset to study multimodal neural attention in visual question answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VQA-MHUG: A gaze dataset to study multimodal neural attention in visual question answering,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.509216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.178967Z digest=sha256:a6297291bb3be64594860c0f1521eb42a12bab76251f35eb2e4c742fff9f3049

Observation 8c632c78-07be-4139-a733-c27e553b0ce6 · outbound

This paper cites Eye gaze tells you where to compute: Gaze-driven efficient vlms,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Eye gaze tells you where to compute: Gaze-driven efficient vlms,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.497438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.182640Z digest=sha256:8e3876d8fbaf26633d68e6c92815567651ae413c74fa79b6b86993e38227e7a9

Observation ea0c5c9b-871a-45d3-924c-b67571e48b83 · outbound

This paper cites Saliency detection: A spectral residual approach,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency detection: A spectral residual approach,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.485099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.186564Z digest=sha256:7274ede6a0a602e66d59d6691d866fd182d3fe21f84caa156b13cbcfcd6c231b

Observation 2984d488-75ea-4a56-b6c8-3cfd71b74027 · outbound

This paper cites Saliency driven perceptual image compression,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency driven perceptual image compression,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.473103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.190484Z digest=sha256:6d82960edf76f01d4e98ea2fdfacfff3f074cf6b32d35e12773ecde11183cc66

Observation 46657777-0be2-4c24-a490-16df0595a874 · outbound

This paper cites Visual cropping improves zero-shot question answering of multimodal large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Visual cropping improves zero-shot question answering of multimodal large language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.461376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.195307Z digest=sha256:b35b8e7f88b46da36443148a78d311f97bd6c8aa6a4c84e616d19b1026ae9eec

Observation 0661ec84-8aed-4a0c-9051-4f3a6796db86 · outbound

This paper cites Edge assisted real-time object detection for mobile augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Edge assisted real-time object detection for mobile augmented reality,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.450028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.199136Z digest=sha256:be1fac3e8f1d5bb70dcdfa7dfea157146d3907ec3c57b11f50983a7d272e0aaa

Observation 6d1ff298-076a-42f4-9024-d3724fdc2cd0 · outbound

This paper cites Glimpse: Continuous, real-time object recognition on mobile devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Glimpse: Continuous, real-time object recognition on mobile devices,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.438645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.202934Z digest=sha256:0f34a493c4e375a3ebb25993212bdef4f6b8834f9db059e9d6ef1afb32d44a3e

Observation bb99a795-38e0-4c1d-834e-d84387269ab0 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.426348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.207129Z digest=sha256:64b98343e46cf35637b19a52b59b062aeb9c6f5cd883b3fed7f19cdda0fa0193

Observation a4b4c0c5-25ef-4e90-8c8c-a85746a2e407 · outbound

This paper cites MMBench: Is your multi- modal model an all-around player?,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MMBench: Is your multi- modal model an all-around player?,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.414074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.211247Z digest=sha256:0f4cda717b4a5ac78ce01d476665a34a7dbba9d74d85addf92c79e4d663beacb

Observation af7e46f8-ecda-49b8-af04-3f441fd59c3e · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.401510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.215889Z digest=sha256:a15217f38d2817ae4c91744a114666ec858b51dd617d14b7e749dda2d6ae6a95

Observation 2659ab9f-0bc5-4bfb-bbd5-c13dde02c5d8 · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.389142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.219803Z digest=sha256:102dd269e8064aeb361c87d75d0bab478f47f6a176c71613622d6f2359d10cab

Observation fab0333c-8dcf-4f0b-a835-2e2417777199 · outbound

This paper cites HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.376588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.223759Z digest=sha256:b0b254f1e0f7172d86c9259ba95e795364ed5c7e0363db5835b782c73533a536

Observation 92215827-6b85-4a43-9454-1c9eaaba8522 · outbound

This paper cites WearVQA: A visual question answering benchmark for wearables in egocentric authentic real-world scenarios,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems WearVQA: A visual question answering benchmark for wearables in egocentric authentic real-world scenarios,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.364074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.228109Z digest=sha256:2ddbbe15927a839df82b2bf64380ccfb71c17ee20ac15816c5936303f35b26a9

Observation ea3fede4-c24d-4cf7-a57f-7641895da963 · outbound

This paper cites SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.232005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.232005Z digest=sha256:cf6a3f0ec2c09650b372b63dc3f361ecff3295ff8871fc618bd975688366731d

Observation 9510abb9-44a5-4dd3-aab4-f6334dcd9568 · outbound

This paper cites CRAG-MM: Multi-modal multi-turn comprehen- sive RAG benchmark,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems CRAG-MM: Multi-modal multi-turn comprehen- sive RAG benchmark,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.238653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.238653Z digest=sha256:d1dcacb1ae4c40f91d700bec35d261c000367a103d9d624216a395aa1c202600

Observation d117e68e-2597-4511-8a8f-cd4b438c1b90 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems OK-VQA: A visual question answering benchmark requiring external knowledge,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.351783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.242494Z digest=sha256:17a7a62df7c795b23b517c39b2c4cc6d1818023613277f18de7f4a2072cac6c9

Observation a41e1861-e52a-4ab1-ba2f-2da8df2db08a · outbound

This paper cites MISAR: A multimodal instructional system with augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MISAR: A multimodal instructional system with augmented reality,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.339738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.246484Z digest=sha256:a71b6227b8e9bc9c826660cbacb0c239eb34c0150ef2e279f5ffc0599d67e793

Observation e10d0e33-8929-4845-b0c5-9b401e07c111 · outbound

This paper cites Guided Reality: Generating visually-enriched AR task guidance with LLMs and vision models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Guided Reality: Generating visually-enriched AR task guidance with LLMs and vision models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.328278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.250850Z digest=sha256:4621fdafd7879bfbb32500a3a884f6433df79273a1b7076981552b89622c9762

Observation 54592885-ffbf-410c-9493-1fea13cc1cd7 · outbound

This paper cites EmBARDiment: an Embodied AI Agent for Productivity in XR.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EmBARDiment: an Embodied AI Agent for Productivity in XR

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.254659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.254659Z digest=sha256:7124b4b0df59d5247b3cebf33f492535ce7d862b49e643da9ad6cc6ee90ef401

Observation ef168b0b-05ee-4c8c-acfc-835114242bdf · outbound

This paper cites GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.315950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.261066Z digest=sha256:e3a652d7c5467fbbd6330947cca560c7914ce7f7771d52af93b5807cb0b24f27

Observation 743d7ba5-8ffe-416b-9a38-98f7ed88fb16 · outbound

This paper cites Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.701778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.265562Z digest=sha256:4ccb04a48bdc688c49185fcdd9dce3ceabc4830d851ff2aeb46957fc17f2a4bc

Observation 01dd9cdc-092c-4bb0-bc95-48c813e8338f · outbound

This paper cites Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.676748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.270631Z digest=sha256:da8f1d47fd541d30c7879d64b363e57e3dab4c9d19d30f62ecf8f0e12e7cb11f

Observation 6676f53c-2d73-4ea9-8c37-0990526deebc · outbound

This paper cites Exploring the use of VLMs for navigation assistance for people with blindness and low vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Exploring the use of VLMs for navigation assistance for people with blindness and low vision,

Reference 62

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:50:19.657100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.275112Z digest=sha256:49e2a9e3942a07e8526dd0261a32d5739be8b1a94b038ea18bbb74101ed1ef46

Observation ed1391d9-df16-421e-b77a-79a1b2965b96 · outbound

This paper cites BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.569534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.280108Z digest=sha256:2df589563e1e313daa1497bf89683c92ae38818b008a91891278d6b690bf6acc

Observation 22753e43-642f-43cd-833a-286ca03f7fa5 · outbound

This paper cites Objective assessment of the webp image coding algorithm,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Objective assessment of the webp image coding algorithm,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.304358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.284581Z digest=sha256:4552d70f106601cde2fc22f7bf6e9ce3e87fa076b9c2d8d8e6e9f33cda7dac1c

Observation 01bace46-f84d-4361-85c5-794780a21227 · outbound

This paper cites Vision - claude api docs,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Vision - claude api docs,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.292984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.288751Z digest=sha256:a3ada34b6f68d5de34cfeb956217da14166bdba5531aa18220932843397c871c

Observation fa6fb7df-483a-4d21-956c-805c2a386197 · outbound

This paper cites Image understanding — gemini api,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Image understanding — gemini api,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.281030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.292821Z digest=sha256:fec5fdacb4edc99d1bc4fb3263096319b4c1be25e2905b3622311b8fdf075aa5

Observation 41af0a65-9a0f-4dd8-8072-cab64b6cb2a6 · outbound

This paper cites Image interpolation and resam- pling,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Image interpolation and resam- pling,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.268208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.297025Z digest=sha256:202c4b146aae3f293d2391294845e951f2109113b3cb67b51811e01e12c66ed0

Observation f99b59f9-f810-4d14-8e34-e9a36a3fd0cd · outbound

This paper cites Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.256188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.301374Z digest=sha256:78796d49e052a832d95d7b94de215a7b2d23b0a8286d09ab0e776e10f2012f6b

Observation 6c8b19a3-96f7-42bb-86d6-9061eba41935 · outbound

This paper cites V oila-a: Aligning vision-language models with user's gaze attention,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems V oila-a: Aligning vision-language models with user's gaze attention,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.242817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.306560Z digest=sha256:9cc14fb717fdd4d5f9b581dc1210803728f031d5f798fb06f718e5f4cd982f85

Observation f87e2bc5-b3fd-4be5-8c95-6ba2989c591d · outbound

This paper cites EgoVQA—an egocentric video question answering benchmark dataset,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EgoVQA—an egocentric video question answering benchmark dataset,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.228121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.310732Z digest=sha256:e6e10d0cff32ae28b18cdbe3886c3a3fd8d1dedf86e897f1398fdfd5979ae180

Observation 9bf1c15d-5ff4-405f-8e17-d64dfc0701e2 · outbound

This paper cites Neural machine translation of rare words with subword units,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Neural machine translation of rare words with subword units,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.216201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.314944Z digest=sha256:6d1dbbf22dc46ba308de0df2aa23a73a00d4eb80ee83754aedbdba59d03801cf

Observation 2fd64ec6-85d1-4967-b5d8-38c431fa085d · outbound

This paper cites What are tokens and how to count them?.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems What are tokens and how to count them?

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.204919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.319134Z digest=sha256:d43bb7269c3e9b85539da600367131175afaa00c3250bda2f1d59220d00045da

Observation 5da132dd-8582-4cbc-8c93-f29ca696bc7f · outbound

This paper cites Saliency based image crop- ping,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency based image crop- ping,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.192706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.323268Z digest=sha256:3363d24ab4bdb446d6614cc37e6e771e5ed432a72efe7b4dc66690313f3c855e

Observation 2794bdd8-29c7-47b4-a08d-47226a09f943 · outbound

This paper cites Benchmarking deep learning models for object detection on edge computing devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Benchmarking deep learning models for object detection on edge computing devices,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.179180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.327140Z digest=sha256:dc34d37ba69cff897e46473ba16ed3644ee7909ee6a4c6e33ece760f157c36a9

Observation 89aca3b1-210a-4f3c-a678-ffd30360e0e6 · outbound

This paper cites Region-of-interest extraction method to increase object-detection performance in remote monitoring system,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Region-of-interest extraction method to increase object-detection performance in remote monitoring system,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.166519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.330909Z digest=sha256:70996d2cbf5d2fcf5e678f0dba7db2550affb30d3358eed15874d7e5f759e8b4

Observation d50df610-97c5-461a-ad07-33e8fbcfb191 · outbound

This paper cites Improving automatic VQA evaluation using large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Improving automatic VQA evaluation using large language models,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.154479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.335462Z digest=sha256:a5bb1e18015480a7f86d36545d164957bd4fa0f85861011f5c6271b493c682da

Observation 33234788-e51a-4184-b466-9825241e5b41 · outbound

This paper cites Mind the uncertainty in human disagreement: Evaluating discrepancies between model predictions and human responses in vqa,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mind the uncertainty in human disagreement: Evaluating discrepancies between model predictions and human responses in vqa,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.142051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.339617Z digest=sha256:8aa64e829fa5b6e830cc4d34eb4381954367fa32c3c01caef3d0b1ef60308aee

Observation 0f4ff4da-137d-44ba-8cf0-faab0953198f · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.128003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.344182Z digest=sha256:ba632902de32ab4180a2f45e9e2e20d223670922ca5d8316c1b8a9d6e6aa3e07

Observation afd1f3e0-03c3-4560-98cf-b74131d5675a · outbound

This paper cites G-Eval: NLG evaluation using GPT-4 with better human alignment,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems G-Eval: NLG evaluation using GPT-4 with better human alignment,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.113194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.348271Z digest=sha256:4501e821e7684f9cce608fb9c4224ebf12198029c7b1236db8d1e781963df7e5

Observation 57d015c5-2614-4081-8fbc-ef1c06409793 · outbound

This paper cites Chatar: Conversation support using large language model and augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Chatar: Conversation support using large language model and augmented reality,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.099459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.352344Z digest=sha256:eafcf3fd9c90a0c550efcbdce2feed6bbc833204303d0c61a50a534216f8ef98

Observation 0db091eb-cd22-4e25-9b2a-5515d36fb4b3 · outbound

This paper cites Next-generation networking and edge computing for mixed reality real-time interactive systems,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Next-generation networking and edge computing for mixed reality real-time interactive systems,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.086379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.356200Z digest=sha256:df0ad1be09ee8ad69a0975368c257c1cd690b6cfe88b0fcd435e070b7500a0e0

Observation 43a2847e-f2a4-4cdc-866c-3ed4e713706d · outbound

This paper cites An egocentric vision-language model based portable real-time smart assistant,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems An egocentric vision-language model based portable real-time smart assistant,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.071213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.360215Z digest=sha256:4e2e52a2cf507fc064c145e8078070f64492e6dbb81a617d1c6342b3d62da308

Observation 4cbd87af-581d-45c3-a2de-fb96d7c5dbfa · outbound

This paper cites Teaching LLMs to see and guide: Context-aware real-time assistance in augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Teaching LLMs to see and guide: Context-aware real-time assistance in augmented reality,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.364187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.364187Z digest=sha256:3353186097faea0b97a31dd2bfd9a0b5355fc7680a5883b355b5305faf25fd05

Observation 3999f9fc-18db-49bc-85a5-767882c1d7fb · outbound

This paper cites Gartner predicts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gartner predicts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.057437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.368841Z digest=sha256:8010a7090f32a346cae4418e843e027c1cd053e96d7119ae41f2b64f2a8d8acf

Observation 9ad2d767-1a29-49bb-8fc9-b5b4f177566a · outbound

This paper cites Deep learning-based object detection in augmented reality: A systematic review,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep learning-based object detection in augmented reality: A systematic review,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.372822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.372822Z digest=sha256:a2891a17ccd5fdd6a8921c7354ee631a4614e2d52c7f5c0eedd884b36c7ecb81

Observation b3c695f9-5789-4416-885c-8d6bf76902b8 · outbound

This paper cites Realtime API with WebRTC,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Realtime API with WebRTC,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.036424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.376497Z digest=sha256:e24c22b2912bf86fdfb544ae5e6a8c5a502e336ffd0feccccbe1623ac6d33ed7

Observation 2437408b-3e1a-4cee-9a4e-18cefaaabc6c · outbound

This paper cites Realtime API with WebSocket,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Realtime API with WebSocket,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.022217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.380300Z digest=sha256:d92640fea9081fd94d7a11b30a5084d18ab55f3ab88ebdabbf92b21887591429

Observation 0b4103ae-d36f-4f62-b6d0-c6c63b65fa68 · outbound

This paper cites Bench- marking the effects of operating system interference on extreme-scale parallel machines,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Bench- marking the effects of operating system interference on extreme-scale parallel machines,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.006413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.384617Z digest=sha256:61dc13db5ddabed67feee4b906e99821cd753e8ceec2219d897450933bebf8ae

Observation 1f047041-fa3f-4727-9554-0090e9e3532e · outbound

This paper cites Understanding the causes of performance variability in HPC workloads,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Understanding the causes of performance variability in HPC workloads,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.990984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.389165Z digest=sha256:cc36e20879eefdd03c452082279b4c0960542eefa03838108fc15afc22ba56ea

Observation 129c06b6-3e73-46a7-b6b0-c9eac695ac01 · outbound

This paper cites Overhead Measurement Noise in Different Runtime Environments.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Overhead Measurement Noise in Different Runtime Environments

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.466260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.392944Z digest=sha256:93231eb2e95d1323abed5fd8c0885657f8c748dfb4f23a63848e6de43eea3646

Observation b5e9adb5-74f6-44ab-8670-c8f301c10ab1 · outbound

This paper cites Using microbenchmarks to evaluate system performance,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Using microbenchmarks to evaluate system performance,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.976627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.397014Z digest=sha256:d880f01b8e01b5d033fd1c474db0a25b1e171adc4e1f94fd67edb67692c2257b

Observation 308695cf-3a57-4dd8-945a-a254dd5502f4 · outbound

This paper cites Beyond inference: Performance analysis of DNN server overheads for computer vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Beyond inference: Performance analysis of DNN server overheads for computer vision,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.962037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.400903Z digest=sha256:d0a7603f1577ce04d3847cba487caaf01dc20af6ed48fbae0ff54f8e963dc2f6

Observation 199f6712-add6-4936-95dd-b41d3f3fd1a6 · outbound

This paper cites EdgeYOLO: An edge-real- time object detector,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EdgeYOLO: An edge-real- time object detector,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.948887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.405290Z digest=sha256:3d4a6ef7575a9d06eac63140d9ba5c890f2a582f0691eaea8865f35ff9f8943b

Observation 464bb5f5-a174-40fb-860b-14cfe70ae457 · outbound

This paper cites Images and vision — calculating image tokens.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Images and vision — calculating image tokens

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.935016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.409358Z digest=sha256:87323fb3888f1ed0b5ddc2f793c0d5595ef163b9b769c873998f43bde2653e82

Observation 7d4d740a-2e86-46f3-b44f-b283bba3fcf4 · outbound

This paper cites Per-part media resolution (Gemini 3 only).

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Per-part media resolution (Gemini 3 only)

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.920442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.413398Z digest=sha256:d27e9b66a09ac84eac650555b771b594c28e357dd7c3fcbd560ad583e61f71f0

Observation 279c4101-1146-4c54-a220-213d08828304 · outbound

This paper cites Gemini 3 Flash — model documentation.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gemini 3 Flash — model documentation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.904265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.417282Z digest=sha256:b4af03d7400639b1f1ead5ffe6ae914aa09fe2cae3d0515868fe5baffd58e0d3

Observation 6e963d5a-b3a1-4dab-97b1-0a1dc045532d · outbound

This paper cites Gemini 3 developer guide — generateContent API.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gemini 3 developer guide — generateContent API

Reference 97

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:50:19.877165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.424861Z digest=sha256:b2d0ccdf4b39047ec5294d844cabd430696f9cd45bc2692ec25f163da435bca4

Observation b0674577-89e0-4df8-88f0-db6abe44eb2d · outbound

This paper cites Accessed: 2026-05-28.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Accessed: 2026-05-28

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.891144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.421183Z digest=sha256:e0c63fb4f3b59a9c0306b2e49684cd29b138ee1f5091a6b0785b425c83111e32

Pith citing papers

No inbound Pith citation observations are available.