Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:43.674756Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2505.23308.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:43.674756Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:38.024169Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:52:44.372937Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e2364817-5c38-4f32-9a3e-9b8f8f5f8f4d · outbound
Spoken question answering for visual queries Spoken question answering for visual queries
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e16c0d6f-5043-429f-a446-3ff63f51e23a · outbound
Spoken question answering for visual queries Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e98c7758-57d8-4dc9-a039-69fc9a872a89 · outbound
Spoken question answering for visual queries Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f7ce5fa8-0945-466d-b639-8338190f1428 · outbound
Spoken question answering for visual queries Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c1e6c72-a674-4861-9221-bfa2c5fb12ac · outbound
Spoken question answering for visual queries Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 32026afa-0ceb-47f6-9fe5-46071c63018c · outbound
Spoken question answering for visual queries Visual question answering (VQA) attempts to describe, locate, and reason regarding some visual input [9, 10, 11]
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d34e2aa7-ba44-486e-8560-b63ff416c30f · outbound
Spoken question answering for visual queries The LLaV A model extends a text-based, generative, large language model (LLM) for visual question answering (VQA) by allow- ing visual information input from images
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b17d589f-2748-428b-a728-039fae1af24f · outbound
Spoken question answering for visual queries Speech-only datasets We pre-train the speech projector using the English subset of Multilingual LibriSpeech (MLS) [29]
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation daf03e84-b161-4e47-9b23-0e1840adce69 · outbound
Spoken question answering for visual queries The speech tower is composed of a trained Whis- per encoder2 and a speech projector
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e1caa170-4215-4589-bb19-ebc6eb4e0e7c · outbound
Spoken question answering for visual queries Spoken VQA Table 1: Performance across VQA benchmarks (StyleTTS2 / F5-TTS)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 402bd894-ee0f-423a-bb1f-48265960ac09 · outbound
Spoken question answering for visual queries In contrast, our SVQA approach maintains more robust performance across all bench- marks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aac2e1eb-c748-44a4-8e75-9009deab37aa · outbound
Spoken question answering for visual queries A significant portion of the effort was spent on building both the speech and SVQA datasets
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dfb14b90-ce57-4463-a04f-2c314bba2a29 · outbound
Spoken question answering for visual queries VQA: Visual question an- swering,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d00fa8da-edb2-4a53-b44e-11a9bdcf7428 · outbound
Spoken question answering for visual queries High-resolution im- age synthesis with latent diffusion models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f26246e6-ec86-4efd-a0b7-f86657cc9403 · outbound
Spoken question answering for visual queries SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca87ffb-0cc6-4a61-848e-5732005e753e · outbound
Spoken question answering for visual queries Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b54c6fbb-5bc7-4da3-bc43-931e77ecb537 · outbound
Spoken question answering for visual queries StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d0b38f-ad84-4dd7-a47f-a1bcbaed0e42 · outbound
Spoken question answering for visual queries Visual instruction tuning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc4346d-99e6-414a-9ab1-51eae3e941cc · outbound
Spoken question answering for visual queries Robust speech recognition via large-scale weak supervision,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d241ecdc-97ca-49f4-8f74-926cf05ecb10 · outbound
Spoken question answering for visual queries Learning transfer- able visual models from natural language supervision,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b1cab28e-55cb-48d0-9876-287a93bd654d · outbound
Spoken question answering for visual queries SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb653b5-e2d5-44a1-bb91-c6310593af24 · outbound
Spoken question answering for visual queries A survey on multimodal large language models,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab244cd-062d-451b-91ab-aaaca92a01da · outbound
Spoken question answering for visual queries DocVQA: A Dataset for VQA on Document Images
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2429511-7ddb-47df-bf4e-720dd644097f · outbound
Spoken question answering for visual queries Improved baselines with visual instruction tuning,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43c74d66-fa83-43ee-8e5b-9c61634e3c33 · outbound
Spoken question answering for visual queries Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a24705-9081-43f9-8feb-ba7381aef310 · outbound
Spoken question answering for visual queries InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db54811b-25ba-4edd-babb-f58e6c2e0749 · outbound
Spoken question answering for visual queries ODSQA: Open-domain spoken question answering dataset,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c53e8183-b2c1-4fd3-9d36-d0a2e942df39 · outbound
Spoken question answering for visual queries Knowledge distillation for im- proved accuracy in spoken question answering,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ba7b729c-8fa8-4793-9af5-3f73d883dccf · outbound
Spoken question answering for visual queries Speech-Based Visual Question Answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9d6b1bd0-aec0-4881-b7d8-a5df58d4d206 · outbound
Spoken question answering for visual queries Speech enabled visual question answering using lstm and cnn with real time image cap- turing for assisting the visually impaired,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d1b5061-1f20-4e3b-b5a6-47840979a600 · outbound
Spoken question answering for visual queries SBVQA 2.0: Robust end-to- end speech-based visual question answering for open-ended ques- tions,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f162e18-2bc9-428d-abbe-8ce339d27fe6 · outbound
Spoken question answering for visual queries Towards mul- tilingual spoken visual question answering system using cross- attention,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c5d0ba00-5fa7-4ee4-9981-303085796e8c · outbound
Spoken question answering for visual queries Worldly wise (WoW)-cross-lingual knowledge fusion for fact- based visual spoken-question answering,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 14fae85c-4120-49c6-870a-95df08d854c9 · outbound
Spoken question answering for visual queries TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b80d85-8099-4cdf-b2c6-39320697d568 · outbound
Spoken question answering for visual queries A VQA: A dataset for audio- visual question answering on videos,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 278c41cd-6374-44c3-a969-53654ded78c9 · outbound
Spoken question answering for visual queries Learning to answer questions in dynamic audio-visual scenarios,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 537a3244-d81f-40bc-babc-72a33d79e69a · outbound
Spoken question answering for visual queries Progressive spatio-temporal percep- tion for audio-visual question answering,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 73056d27-3b80-4814-b56a-257c3235605d · outbound
Spoken question answering for visual queries TVLT: Textless vision- language transformer,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 515ba5d2-a398-4411-8afa-99b8b4d4ddc1 · outbound
Spoken question answering for visual queries Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b9aedd00-dd34-4376-aa4c-1946c8e6ac56 · outbound
Spoken question answering for visual queries Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56bcec15-253a-4c0d-820d-17811f98918c · outbound
Spoken question answering for visual queries MLS: A large-scale multilingual dataset for speech research,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f591dbd1-ebf9-4229-94fb-c5ccde3f1f31 · outbound
Spoken question answering for visual queries Audiochatllama: Towards general- purpose speech abilities for llms,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cd6a42c2-5666-4106-8643-858b0e82fd13 · outbound
Spoken question answering for visual queries What matters when building vision-language models?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a796c9b-cd70-41d0-8014-5d5d7a4ca622 · outbound
Spoken question answering for visual queries Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca0c7e0d-86b8-4faf-bfbd-f43cdbcb87f0 · outbound
Spoken question answering for visual queries Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a79f3414-e36a-4df3-ac53-d17663845828 · outbound
Spoken question answering for visual queries HeySQuAD: A spoken question answering dataset,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7bca26dc-174e-4c7e-b9df-abc2fbcf768b · outbound
Spoken question answering for visual queries F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1de9b23-f463-4c5c-a580-d41614ee40de · outbound
Spoken question answering for visual queries LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9995415-57b8-4887-aa51-105d8f1371d9 · outbound
Spoken question answering for visual queries LMMs-Eval: Accelerating the development of large multimoal models,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8ac98437-f6cd-486a-b929-107ca812e6fe · outbound
Spoken question answering for visual queries Your general opinion on the paper should be positive
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2364817-5c38-4f32-9a3e-9b8f8f5f8f4d · inbound
Spoken question answering for visual queries Spoken question answering for visual queries
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.