Pith. sign in

Paper Citation Record · LEDGER

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2305.15255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15255 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:02.306051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:12.799615Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca56d13a-6ae1-47f2-8bb2-cbfae82600fe · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:02.306051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:02.306051Z digest=sha256:d6318b682406e1c7e22ccdddff03dbb25a62e62e4c5e549477733d94d1b63423

Observation a9c77d13-2b1c-4e61-8113-7d64e9720542 · inbound

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering cites this paper.

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:01.447743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:01.447743Z digest=sha256:b5a501b7028e33b893067d8b9518cfdf260c7910ceb0a233d6321b3ab5078e3c

Observation 578e6321-953e-4e70-9458-bd1ed3843068 · inbound

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models cites this paper.

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:32.437237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:32.437237Z digest=sha256:d89bc72165789fd6562cf003e9d9072ec4f5a878c7dea606dcce9e05f2124c2a

Observation 75dea8af-703c-4fbc-ad56-db0becf0e6c3 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.901004Z digest=sha256:1d38f36d5493b2e478c782bdd5c02726d061495bf5f9edf422efc8aabd18fc7a

Observation a82b8b51-06c7-4edb-8cc2-9f5c41cb3203 · inbound

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity cites this paper.

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:29.500611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:29.500611Z digest=sha256:e9184c45a1eab38933777380ef3b2d33fc6f26211608668c7750f13317ab7796

Observation 5d30e757-ba22-422a-b20a-df145a1b26f0 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.440358Z digest=sha256:b8b3df98bedb2047cb38a4ef6e5caa3e2219a90fe22ead3bdddc1ae27924984d

Observation 0c6d9238-9845-4a4d-85db-f2abf53711b6 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.523554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:781d872af78b3471ff1bf6ef9ea9bb5329d20c503ca910523866d938112b81f2

Observation 0fdeb41d-e64d-44fc-b20b-615f62936534 · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.808724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.808724Z digest=sha256:89c8fc047c219ab2c50acd213ecff357b796f62e8b079f9f535a6b32c140857a

Observation a9cf42d6-53ed-4d9f-9519-77c0fc8aab83 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:35:18.336465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:c3c7867a9ef1dff4e05b6a3a31c3e6a826bb6b52e31987a97267ac9cea15a45b

Observation e36d694d-e081-4521-af19-311c41dca4b8 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:44:07.731649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:083f244ce36e49ef86d2eda590c591bdb5e4763efbc21b3d14c5ac0dbfa5696d

Observation f41ea930-1e84-4425-8d15-f62eaa6ab6c1 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:3fd174d198aedf40b5d1946ff5bfb88a232a2ae9b479729ecc2f6ad3f8b9f66c

Observation c9855641-f6f8-4435-b6ac-d3fd4a17243b · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:56.112480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:23bb3a15a45e846bf075f529aeec8e135351206ded2b0dd568b812d16ac9adbc

Observation dd0c267f-3395-41ff-8ed8-11076ca15720 · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 46

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T18:43:50.740325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:d0b79d195598f8c13c9baf174d9cd154da9b946af710a420e4379170aa0a0c8d

Observation 5b065715-e004-4abb-b17d-6528a168b85c · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.399346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:e19f0bbff229403dc5b443e87fa047b2cea0d42388d49818e92cc03e1cab1da8

Observation 814d8f26-395c-4231-92d2-c67a346ae863 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.801146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:ba51b97a078cf041a62bf45be05819ebbf8b305675b8d6b29e28877fed8f00a5

Observation 8974942b-412e-4800-887f-6fa99e2acc01 · inbound

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars cites this paper.

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.396461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:38:28.718164Z digest=sha256:f31d65592d3b1253cdc4bc5fd95372fdd1cd00e8ec48d401a390e8a263e267a7

Observation bee2084b-cdfb-4f8b-9679-bb172797f891 · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:17.845680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:17.845680Z digest=sha256:6478ade2380603614fa9748795bb4a4aa687509e2f749b4f8851fd5f76d42626