Pith. sign in

Paper Citation Record · LEDGER

Spoken question answering for visual queries

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2505.23308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23308 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:43.674756Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:38.024169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:52:44.372937Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2364817-5c38-4f32-9a3e-9b8f8f5f8f4d · outbound

This paper cites Spoken question answering for visual queries.

Spoken question answering for visual queries Spoken question answering for visual queries

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:52:44.527282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.024169Z digest=sha256:836140120ced606b4fcd1d49b98fba166704bbf20a24b127894764a64ab3354e

Observation e16c0d6f-5043-429f-a446-3ff63f51e23a · outbound

This paper cites an unresolved cited work.

Spoken question answering for visual queries Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:52:51.843899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.089468Z digest=sha256:a4dd53c2f1f2feea2075dd5b110242fcd4c9e1688c13f5eacc2deb174205d9b1

Observation e98c7758-57d8-4dc9-a039-69fc9a872a89 · outbound

This paper cites an unresolved cited work.

Spoken question answering for visual queries Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:52:51.668743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.157075Z digest=sha256:33f77a4409e6747ed1a3c95503cd618a84cf1414fb69cd94ad32e21f2954a10c

Observation f7ce5fa8-0945-466d-b639-8338190f1428 · outbound

This paper cites an unresolved cited work.

Spoken question answering for visual queries Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:52:51.427848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.272015Z digest=sha256:9b867c64fedc7282737c45c42ff3b11c5f8d8defaf155027b485b5dde73ff60b

Observation 6c1e6c72-a674-4861-9221-bfa2c5fb12ac · outbound

This paper cites an unresolved cited work.

Spoken question answering for visual queries Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:52:51.237876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.368166Z digest=sha256:828e7edf4432912fcd6eca5af0ef2eefe3ab5d9c3e02bb0ef4b9fe81aeacc735

Observation 32026afa-0ceb-47f6-9fe5-46071c63018c · outbound

This paper cites Visual question answering (VQA) attempts to describe, locate, and reason regarding some visual input [9, 10, 11].

Spoken question answering for visual queries Visual question answering (VQA) attempts to describe, locate, and reason regarding some visual input [9, 10, 11]

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:51.107308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.452799Z digest=sha256:4e37b63dc938f65699974e33d2508c130d7fcedeb98827b49ea0f5827ff5cc52

Observation d34e2aa7-ba44-486e-8560-b63ff416c30f · outbound

This paper cites The LLaV A model extends a text-based, generative, large language model (LLM) for visual question answering (VQA) by allow- ing visual information input from images.

Spoken question answering for visual queries The LLaV A model extends a text-based, generative, large language model (LLM) for visual question answering (VQA) by allow- ing visual information input from images

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:50.892877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.531846Z digest=sha256:381fe8f3e1593c62e36c5c07560a8a5b1a5efc8b36031032d9a4931eb7fcd25a

Observation b17d589f-2748-428b-a728-039fae1af24f · outbound

This paper cites Speech-only datasets We pre-train the speech projector using the English subset of Multilingual LibriSpeech (MLS) [29].

Spoken question answering for visual queries Speech-only datasets We pre-train the speech projector using the English subset of Multilingual LibriSpeech (MLS) [29]

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:50.734359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.603024Z digest=sha256:4515599144d7d61810171d79de99a818ef065e3af415292d37de2977d08f6c6c

Observation daf03e84-b161-4e47-9b23-0e1840adce69 · outbound

This paper cites The speech tower is composed of a trained Whis- per encoder2 and a speech projector.

Spoken question answering for visual queries The speech tower is composed of a trained Whis- per encoder2 and a speech projector

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:50.528588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.689074Z digest=sha256:b36b54c4df096f9523cb8280631a25e2aed31ca9092eeb96cec2fa3a2b0ba909

Observation e1caa170-4215-4589-bb19-ebc6eb4e0e7c · outbound

This paper cites Spoken VQA Table 1: Performance across VQA benchmarks (StyleTTS2 / F5-TTS).

Spoken question answering for visual queries Spoken VQA Table 1: Performance across VQA benchmarks (StyleTTS2 / F5-TTS)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:50.360688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.759945Z digest=sha256:aea9ba8e7f6b3ad078d7a14bcec0fc8e0c4bdf34abf81eb03269e864dcd203fb

Observation 402bd894-ee0f-423a-bb1f-48265960ac09 · outbound

This paper cites In contrast, our SVQA approach maintains more robust performance across all bench- marks.

Spoken question answering for visual queries In contrast, our SVQA approach maintains more robust performance across all bench- marks

Reference 11

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:52:50.146588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.839356Z digest=sha256:09e4cc10c2ac322d7780148083b0dd37c863ab7e9f96ce2270d02ae180b50d46

Observation aac2e1eb-c748-44a4-8e75-9009deab37aa · outbound

This paper cites A significant portion of the effort was spent on building both the speech and SVQA datasets.

Spoken question answering for visual queries A significant portion of the effort was spent on building both the speech and SVQA datasets

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:49.940924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.904415Z digest=sha256:0f65bdd680aea412ca5ba3f2f41c425c1fe89a5652981b48b33e980a2e0a9a61

Observation dfb14b90-ce57-4463-a04f-2c314bba2a29 · outbound

This paper cites VQA: Visual question an- swering,.

Spoken question answering for visual queries VQA: Visual question an- swering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:49.760777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:39.016333Z digest=sha256:5e8d2b931ef114aaf5f95694c50dce49e2214d23ab2eb0512167f5fea924a9b4

Observation d00fa8da-edb2-4a53-b44e-11a9bdcf7428 · outbound

This paper cites High-resolution im- age synthesis with latent diffusion models,.

Spoken question answering for visual queries High-resolution im- age synthesis with latent diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:49.519865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:39.103424Z digest=sha256:477e8f9dc7430063f776b7115fe0cf7d0f0c69a0940e44f9454b477d0b83ff09

Observation f26246e6-ec86-4efd-a0b7-f86657cc9403 · outbound

This paper cites SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering.

Spoken question answering for visual queries SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.184008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.184008Z digest=sha256:9d7c162aaf320f49ac2c0101dd6b628f257b7955071bcb12acc7231e03bad994

Observation dca87ffb-0cc6-4a61-848e-5732005e753e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Spoken question answering for visual queries Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.261148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.261148Z digest=sha256:455fdc3e8f1610e43af85007973b2f5fcdeb94f8deedb84e7b88db9f8adf949c

Observation b54c6fbb-5bc7-4da3-bc43-931e77ecb537 · outbound

This paper cites StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.

Spoken question answering for visual queries StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.339674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.339674Z digest=sha256:367bfc311e7d60dc0df3fa99fef4248635729aa003a5c7cab126d1bafd3b5495

Observation 66d0b38f-ad84-4dd7-a47f-a1bcbaed0e42 · outbound

This paper cites Visual instruction tuning,.

Spoken question answering for visual queries Visual instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.434164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.434164Z digest=sha256:426ecbcafb5bc5875aaa3abb674839a526b5758420df88dc27448b3cfca6dc02

Observation 9dc4346d-99e6-414a-9ab1-51eae3e941cc · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Spoken question answering for visual queries Robust speech recognition via large-scale weak supervision,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:49.312406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:39.518167Z digest=sha256:887d2f529ab329f85b835d06660eb7259daf8f3ab5ffa2df93953db60e90c229

Observation d241ecdc-97ca-49f4-8f74-926cf05ecb10 · outbound

This paper cites Learning transfer- able visual models from natural language supervision,.

Spoken question answering for visual queries Learning transfer- able visual models from natural language supervision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:49.088898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:39.600112Z digest=sha256:e842442b3581e0c8971e4425cbcf05a954754fed77b4c9545044ebfe6475b746

Observation b1cab28e-55cb-48d0-9876-287a93bd654d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Spoken question answering for visual queries SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.686996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.686996Z digest=sha256:f681ceee7ba422af27b37595cc07315112b11d3c7ba1db712ea86927923773f4

Observation 0fb653b5-e2d5-44a1-bb91-c6310593af24 · outbound

This paper cites A survey on multimodal large language models,.

Spoken question answering for visual queries A survey on multimodal large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.815338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.815338Z digest=sha256:4d35fff0505cfad666fbff70a4969e98cd2e3ffdb75869bfc768da3c58d8313f

Observation eab244cd-062d-451b-91ab-aaaca92a01da · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Spoken question answering for visual queries DocVQA: A Dataset for VQA on Document Images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.942212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.942212Z digest=sha256:156df7b65921d17c33df28ede8ab435c2944e5d9f415d2e5783a91de656515b3

Observation e2429511-7ddb-47df-bf4e-720dd644097f · outbound

This paper cites Improved baselines with visual instruction tuning,.

Spoken question answering for visual queries Improved baselines with visual instruction tuning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:48.899988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:40.084560Z digest=sha256:eed902336d79e84cefa86d1a2df11ce986c4e7e7e0d637d8774af0fe52db5929

Observation 43c74d66-fa83-43ee-8e5b-9c61634e3c33 · outbound

This paper cites Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence.

Spoken question answering for visual queries Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:40.237406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:40.237406Z digest=sha256:ce1952dfaf6c1f8b1a5c69e2c83bbe84f15fc7e70be718668e3bd19cdb3c5b3a

Observation d6a24705-9081-43f9-8feb-ba7381aef310 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Spoken question answering for visual queries InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:40.386603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:40.386603Z digest=sha256:5f3609e4e63cb1fa72da231b4ada0f44d5d9c5a4d3958e0c53e221f58014a2d9

Observation db54811b-25ba-4edd-babb-f58e6c2e0749 · outbound

This paper cites ODSQA: Open-domain spoken question answering dataset,.

Spoken question answering for visual queries ODSQA: Open-domain spoken question answering dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:48.659526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:40.471458Z digest=sha256:c22f871b117d571f122103c7ae0a9de0c70d7b930914e461fef2f5a408174ad2

Observation c53e8183-b2c1-4fd3-9d36-d0a2e942df39 · outbound

This paper cites Knowledge distillation for im- proved accuracy in spoken question answering,.

Spoken question answering for visual queries Knowledge distillation for im- proved accuracy in spoken question answering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:48.399122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:40.598190Z digest=sha256:d08cb5b6d69f6120d321325307ef78702169bf1ccb1e185e67585f0fdc29a71d

Observation ba7b729c-8fa8-4793-9af5-3f73d883dccf · outbound

This paper cites Speech-Based Visual Question Answering.

Spoken question answering for visual queries Speech-Based Visual Question Answering

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:52:44.095480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:40.748552Z digest=sha256:12f4b8b27039402c0e2196a5624c4f864a03e1a8381e203504d91aba1c99b309

Observation 9d6b1bd0-aec0-4881-b7d8-a5df58d4d206 · outbound

This paper cites Speech enabled visual question answering using lstm and cnn with real time image cap- turing for assisting the visually impaired,.

Spoken question answering for visual queries Speech enabled visual question answering using lstm and cnn with real time image cap- turing for assisting the visually impaired,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:48.150661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:40.913421Z digest=sha256:f89029b5c56e4153f04fd21de4297149454039daac9353130a94e7aa21e06f90

Observation 0d1b5061-1f20-4e3b-b5a6-47840979a600 · outbound

This paper cites SBVQA 2.0: Robust end-to- end speech-based visual question answering for open-ended ques- tions,.

Spoken question answering for visual queries SBVQA 2.0: Robust end-to- end speech-based visual question answering for open-ended ques- tions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:48.009869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.044499Z digest=sha256:1be59c36bdd0b8afc330dc3bc75768db2bcbd0c3cd238492247964e523212e14

Observation 1f162e18-2bc9-428d-abbe-8ce339d27fe6 · outbound

This paper cites Towards mul- tilingual spoken visual question answering system using cross- attention,.

Spoken question answering for visual queries Towards mul- tilingual spoken visual question answering system using cross- attention,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:47.807587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.156832Z digest=sha256:8f7287155b44e404dc9dd1c97e0c8ff4fa863b58d0f5e4489f3a32b7d9aac54e

Observation c5d0ba00-5fa7-4ee4-9981-303085796e8c · outbound

This paper cites Worldly wise (WoW)-cross-lingual knowledge fusion for fact- based visual spoken-question answering,.

Spoken question answering for visual queries Worldly wise (WoW)-cross-lingual knowledge fusion for fact- based visual spoken-question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:47.550425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.243290Z digest=sha256:0a91aa716d8903830706fef3532c91a109c7388b18c6286de0529f49ddab00d4

Observation 14fae85c-4120-49c6-870a-95df08d854c9 · outbound

This paper cites TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering.

Spoken question answering for visual queries TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:41.348522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:41.348522Z digest=sha256:215e3207cfec8c0d9c2558388ded851c0b84dc8c58effd7d0b64b3541688ae16

Observation c0b80d85-8099-4cdf-b2c6-39320697d568 · outbound

This paper cites A VQA: A dataset for audio- visual question answering on videos,.

Spoken question answering for visual queries A VQA: A dataset for audio- visual question answering on videos,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:47.291482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.478609Z digest=sha256:7bdfee5e89fced36e30a1a6098a77a3a3a5d19f2bbfb767105507536e268ac92

Observation 278c41cd-6374-44c3-a969-53654ded78c9 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios,.

Spoken question answering for visual queries Learning to answer questions in dynamic audio-visual scenarios,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:46.993375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.621202Z digest=sha256:e66f4a791261f0b4f82b5835fa41311ca105cb5f58c19c920818af3a4d4b8874

Observation 537a3244-d81f-40bc-babc-72a33d79e69a · outbound

This paper cites Progressive spatio-temporal percep- tion for audio-visual question answering,.

Spoken question answering for visual queries Progressive spatio-temporal percep- tion for audio-visual question answering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:46.750107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.738830Z digest=sha256:bcdf20a9b565a78f5951965b4d1501e7c998701b673e1064755300c64aefc23c

Observation 73056d27-3b80-4814-b56a-257c3235605d · outbound

This paper cites TVLT: Textless vision- language transformer,.

Spoken question answering for visual queries TVLT: Textless vision- language transformer,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:46.445785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.863793Z digest=sha256:51dabe42c55f89e36a1df85fbf6ac091bcb5db49115da3009a141f3d9d16c55f

Observation 515ba5d2-a398-4411-8afa-99b8b4d4ddc1 · outbound

This paper cites Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,.

Spoken question answering for visual queries Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:46.101046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:41.936684Z digest=sha256:b583c7e323daad87ca44efacfc51a06081cb9a5245e2bfe05935a0ac96fc2c23

Observation b9aedd00-dd34-4376-aa4c-1946c8e6ac56 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Spoken question answering for visual queries Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:42.077607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:42.077607Z digest=sha256:23be5e1e83efdd2a2aa0c59b76e1fc3a5c3180b4662afa1dfb494d33fbacc079

Observation 56bcec15-253a-4c0d-820d-17811f98918c · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research,.

Spoken question answering for visual queries MLS: A large-scale multilingual dataset for speech research,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:42.260823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:42.260823Z digest=sha256:2ab391c4e41091ed865eb464125df64b016b7077e9ec19d693376b492ca4005b

Observation f591dbd1-ebf9-4229-94fb-c5ccde3f1f31 · outbound

This paper cites Audiochatllama: Towards general- purpose speech abilities for llms,.

Spoken question answering for visual queries Audiochatllama: Towards general- purpose speech abilities for llms,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:45.867320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:42.415832Z digest=sha256:fe630426d0393f1526a47e8ed97c5ae8ffc0272fa07209364bbbf88dc8a4e6e2

Observation cd6a42c2-5666-4106-8643-858b0e82fd13 · outbound

This paper cites What matters when building vision-language models?.

Spoken question answering for visual queries What matters when building vision-language models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:42.565008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:42.565008Z digest=sha256:96eaffc33013ea1a1f048282b817a27cea14719764493555f6b2685fdc581bb2

Observation 6a796c9b-cd70-41d0-8014-5d5d7a4ca622 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Spoken question answering for visual queries Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:42.746933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:42.746933Z digest=sha256:57700a86398f7af7867dce704c57384f8c74b5ab5fbaa1650e24955e5a7bb893

Observation ca0c7e0d-86b8-4faf-bfbd-f43cdbcb87f0 · outbound

This paper cites Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,.

Spoken question answering for visual queries Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:45.627436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:42.904080Z digest=sha256:72fd7f9d08f45706d01206186eda2d38bf577d0d7a26cdd56d0683ac553b3c86

Observation a79f3414-e36a-4df3-ac53-d17663845828 · outbound

This paper cites HeySQuAD: A spoken question answering dataset,.

Spoken question answering for visual queries HeySQuAD: A spoken question answering dataset,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:45.323528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:43.080957Z digest=sha256:bcb6058e661b240c8dbf406a7f5fde309d1fe67a73fcafb4a03736eabb5d5a53

Observation 7bca26dc-174e-4c7e-b9df-abc2fbcf768b · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Spoken question answering for visual queries F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:43.292481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:43.292481Z digest=sha256:f4b60ca424b738cf4cf8b7bd1a0a3a3cf15ccfad37b9b63b4e59faa5b90be30e

Observation c1de9b23-f463-4c5c-a580-d41614ee40de · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Spoken question answering for visual queries LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:43.377377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:43.377377Z digest=sha256:04b2eb744c082f6e52cf9a68ca1e24add5220c6508650307b88cc9555be7b612

Observation c9995415-57b8-4887-aa51-105d8f1371d9 · outbound

This paper cites LMMs-Eval: Accelerating the development of large multimoal models,.

Spoken question answering for visual queries LMMs-Eval: Accelerating the development of large multimoal models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:45.093170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:43.503710Z digest=sha256:7d8fbaf0565653243ffd8d47b88ac9f328c7765f94f065acd9d13e18fd38eb0f

Observation 8ac98437-f6cd-486a-b929-107ca812e6fe · outbound

This paper cites Your general opinion on the paper should be positive.

Spoken question answering for visual queries Your general opinion on the paper should be positive

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:52:44.830579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:43.674756Z digest=sha256:5ce95bda61637dc0e2ca345730debfbe4cfd458e6a1cc7f62e81946fe24cb4f9

Pith citing papers

Observation e2364817-5c38-4f32-9a3e-9b8f8f5f8f4d · inbound

Spoken question answering for visual queries cites this paper.

Spoken question answering for visual queries Spoken question answering for visual queries

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:52:44.527282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:52:38.024169Z digest=sha256:836140120ced606b4fcd1d49b98fba166704bbf20a24b127894764a64ab3354e