Pith. sign in

Paper Citation Record · LEDGER

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

As of 21 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2509.06010.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06010 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:43:20.831117Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:19.280108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:50:19.564824Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a71d3b19-4d05-44c9-bd8a-678cf942a4cf · outbound

This paper cites Vision-language model-based polyformer for recognizing visual ques- tions with multiple answer groundings.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vision-language model-based polyformer for recognizing visual ques- tions with multiple answer groundings

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.704356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:18.629579Z digest=sha256:59089d137bae1cb5d24ba43f5a7eb2f56e36d2c879dc6756194510415f5a3203

Observation 7969617d-6b83-4407-b0c4-90192c9a1b68 · outbound

This paper cites Remote assistance for blind users in daily life: A survey about be my eyes.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Remote assistance for blind users in daily life: A survey about be my eyes

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.597271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:18.722461Z digest=sha256:4aed886c0b8956c1c26cecca30ddd1a577f79c02bb7788ffe327ed0c9cd3a8f3

Observation 038400d4-9bfc-4000-91bf-9a075d73091f · outbound

This paper cites Vqa therapy: Exploring answer differences by visually grounding answers.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vqa therapy: Exploring answer differences by visually grounding answers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.481862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:18.812578Z digest=sha256:0a0bd714bfa11ba00177033114effadc7484a2cea86039863212c6d78fb1124f

Observation 88fb4172-b046-4b0a-97de-69bb7b728e07 · outbound

This paper cites Refining pseudo labeling via multi- granularity confidence alignment for unsupervised cross domain object detection.IEEE Transactions on Image Processing, 2025.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Refining pseudo labeling via multi- granularity confidence alignment for unsupervised cross domain object detection.IEEE Transactions on Image Processing, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.318942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:18.934653Z digest=sha256:dc2a5ec292e86622b984a5cf14c91155d3e43cb26881d58787b67d6f000e2883

Observation 17b92c3e-06ea-40ea-81ea-36dc0b1b2806 · outbound

This paper cites Vqask: a multimodal android gpt- based application to help blind users visualize pictures.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vqask: a multimodal android gpt- based application to help blind users visualize pictures

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.155782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.079205Z digest=sha256:77ccb88b3d7d4e5545df0c07d0a5988ec92544e9565bcb46284fe1e272657b56

Observation c35eb609-729a-4598-892a-f983135fa42f · outbound

This paper cites Towards understanding the use of mllm-enabled applications for visual interpretation by blind and low vision people.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Towards understanding the use of mllm-enabled applications for visual interpretation by blind and low vision people

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.013945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.198124Z digest=sha256:2707503d85cc8abb2e5f2e7b835571663a441f114cdd23f15a9a90dbc23c3046

Observation 61ae3016-6ceb-4779-84ac-d3683cd7abef · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vizwiz grand challenge: Answering visual questions from blind people

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.856931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.299838Z digest=sha256:599a300b45a7f9ba7db2dd27b97837dceda16f7fe3e470d6cd7b6a84164a4038

Observation 74eaca94-984d-4733-a645-be52a4ed3189 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:43:19.427202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:43:19.427202Z digest=sha256:f5d42b55104853381ddd7ac4962df0f771beab1943f351b8ff959b5a0f593965

Observation 22e64f2a-a691-47bb-921b-5355012f4667 · outbound

This paper cites Consistency and uncertainty: Identifying unre- liable responses from black-box vision-language models for selective visual question answering.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Consistency and uncertainty: Identifying unre- liable responses from black-box vision-language models for selective visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.658833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.511234Z digest=sha256:b6e3b9ef9993d866a6d631e357e717600dc5cf1733e585d33464357e6bf05e19

Observation 3e5d52a3-e44c-4d03-858f-be4c7f564a44 · outbound

This paper cites Dual-branch fusion with style modulation for cross-domain few-shot semantic segmentation.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Dual-branch fusion with style modulation for cross-domain few-shot semantic segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.540118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.621669Z digest=sha256:7d142edc6262ef6603ab51d834c39ed49f1833b1f3b27926dd5267cc0031af6c

Observation 13112c59-87cd-4b19-bb27-628a24cc6a46 · outbound

This paper cites Natural language understanding and inference with mllm in visual question answering: A survey.ACM Computing Surveys, 57(8):1–36, 2025.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Natural language understanding and inference with mllm in visual question answering: A survey.ACM Computing Surveys, 57(8):1–36, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.407626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.745316Z digest=sha256:4709fee25805f64d19eb5f37f1426451a8b4fc1905ad4ccf2c4c3c6e3ad9e0cf

Observation 8a0b82b8-71f9-4ab7-9d16-3712a3cb19fa · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.242938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:19.846109Z digest=sha256:3e12391ad7a0f6864145e197e74473a43caa074182c710511f0130c744b1a0a6

Observation 59ab5756-76b5-4888-b221-fabfbf80582e · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Polyformer: Referring image segmentation as sequential polygon generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.135019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.005306Z digest=sha256:86e879dfc0d70ae125d06726ba874f398f54cf5c001dd66a02a7fd15d418aa16

Observation 4cc3c410-d558-428a-a2d3-d7f83cd29085 · outbound

This paper cites An astute assistive device for mobility and object recognition for visually impaired people.IEEE Transactions on Human-Machine Systems, 49(5):449–460, 2019.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users An astute assistive device for mobility and object recognition for visually impaired people.IEEE Transactions on Human-Machine Systems, 49(5):449–460, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.046250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.210949Z digest=sha256:6ddc6165c15ec33bf021e9f14464670b434f302ccf1e53c34856e08c387b75dc

Observation 51efa02b-fb72-4e61-b604-ec675b7df3fa · outbound

This paper cites Dynamic conceptional con- trastive learning for generalized category discovery.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Dynamic conceptional con- trastive learning for generalized category discovery

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.954258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.357680Z digest=sha256:bf8988e8cc141cead2320f3be7d21966af6143432d3ab809dc9798f4e207a0c1

Observation 73ebf983-0b0a-4f26-9844-b897620e2f23 · outbound

This paper cites Advances in few-shot action recognition: A comprehensive review.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Advances in few-shot action recognition: A comprehensive review

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.873693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.466310Z digest=sha256:245fa67f10cac5e3aa77fec4a6c14d7df7ffa80b760bf6d77838c22d92435dc9

Observation 2743f3ba-f3bf-4408-b218-19313375d57a · outbound

This paper cites DARE: Diverse Visual Question Answering with Robustness Evaluation.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users DARE: Diverse Visual Question Answering with Robustness Evaluation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:43:21.034136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.547086Z digest=sha256:7ebb5d0c45ce08e1c73a74bafb0f3653e60fc508538691757cd0f00e58b53e2a

Observation 7cfad7f6-621b-40f6-80f5-cb3990f11da5 · outbound

This paper cites the smart vision glasses.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users the smart vision glasses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.753393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.632394Z digest=sha256:16afaa51fd4fb35a1303e6b8b1b92da8b383e18b8aa5be2674c21f212dd49d4a

Observation 55a58566-607d-42f3-8275-e293688f50dc · outbound

This paper cites A survey of 17 indoor travel assistance systems for blind and visually impaired people.IEEE Transactions on Human-Machine Systems, 52(1):134–148, 2021.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users A survey of 17 indoor travel assistance systems for blind and visually impaired people.IEEE Transactions on Human-Machine Systems, 52(1):134–148, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.627116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.706463Z digest=sha256:3356e3ae9d8bd73db3195bb532ab697cc5f972f3262236ab8cfe09f137dcdea4

Observation 3793b5f3-9b50-4591-b94c-5f3bf8166f85 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.422605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.778291Z digest=sha256:a3aef30edfe2fd905d2a3e49e9dcf7c8b3343f1f40aee1115677acd1d4c9b5c5

Observation c3c93b60-2266-480f-a13a-eabe623cbcae · outbound

This paper cites A survey on vqa: Datasets and approaches.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users A survey on vqa: Datasets and approaches

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.238294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T04:43:20.831117Z digest=sha256:0770061c4e81a1d2784e9f556f035fcbf6519a27931a5684a139032d305e8c38

Pith citing papers

Observation ed1391d9-df16-421e-b77a-79a1b2965b96 · inbound

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems cites this paper.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.569534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:50:19.280108Z digest=sha256:2df589563e1e313daa1497bf89683c92ae38818b008a91891278d6b690bf6acc