Pith. sign in

Paper Citation Record · LEDGER

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2305.15255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15255 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:02.306051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:12.799615Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca56d13a-6ae1-47f2-8bb2-cbfae82600fe · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:02.306051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:02.306051Z digest=sha256:3dc0fcd60d0a001fc545b272e5b53b17511af15ea2066852cbc960d7fe14d928

Observation a9c77d13-2b1c-4e61-8113-7d64e9720542 · inbound

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering cites this paper.

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:01.447743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:01.447743Z digest=sha256:5e28942acbd1a460cb4bec94d07da041d86e475be34cd67e502928a38164e329

Observation 578e6321-953e-4e70-9458-bd1ed3843068 · inbound

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models cites this paper.

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:32.437237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:32.437237Z digest=sha256:381a96a0904ba6fdbad82de82184ae4e87b2764ab7521093026f6562324be921

Observation 75dea8af-703c-4fbc-ad56-db0becf0e6c3 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.901004Z digest=sha256:7f3022f3da7d1789cd6e19a139069e8011bdbd10fe8ce296df9524a2f5502a1f

Observation a82b8b51-06c7-4edb-8cc2-9f5c41cb3203 · inbound

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity cites this paper.

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:29.500611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:29.500611Z digest=sha256:e6c3633dec78eab0b5b7b5bbcc1e167e18b085c5cb552a37b00ddb3abbb9b8af

Observation 5d30e757-ba22-422a-b20a-df145a1b26f0 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.440358Z digest=sha256:f0df1bb5f21891e7591d5c77364090785610ed6add297adf37d4005b2f7c6d0e

Observation 0c6d9238-9845-4a4d-85db-f2abf53711b6 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.523554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:557f704920917bd282076da00e6b86d82844f968cb8247b110b73ff3d35e6506

Observation 0fdeb41d-e64d-44fc-b20b-615f62936534 · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.808724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.808724Z digest=sha256:91e0a8c311a0ec1645d06166c1febffd4f50f916fc8eac0ec2cd2363662902e5

Observation a9cf42d6-53ed-4d9f-9519-77c0fc8aab83 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:35:18.336465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:6e7448a4a0b88f992d0f199fe474f930c07155ba060daafaee89176e06358d49

Observation e36d694d-e081-4521-af19-311c41dca4b8 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:44:07.731649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:378b515f0e65c6005c8c6c401d77c30d27099769642fee67884635a1b61548cb

Observation f41ea930-1e84-4425-8d15-f62eaa6ab6c1 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:22a274d67c2ea3e9f9c4c050ce3ea4dd14edd731fc0cb078b5ccdda6961bd642

Observation c9855641-f6f8-4435-b6ac-d3fd4a17243b · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:56.112480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:e608506bdb495bd88e697343a76232d341ab46dcf136d96b9e9c818ec6ca5293

Observation dd0c267f-3395-41ff-8ed8-11076ca15720 · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 46

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T18:43:50.740325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:172baeaf0ce73deba05f2f352fa4f132490314e9022733214827e225fe3eda8f

Observation 5b065715-e004-4abb-b17d-6528a168b85c · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.399346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:7e98ffe7c6b893817084eb87358ae1a4fc1b46af24634045a86c5f4f2bb847f4

Observation 814d8f26-395c-4231-92d2-c67a346ae863 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.801146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:9d491f48b9a6514f93314b6e4218198843a0ce2e1d12a503791cac3807cddcbc

Observation 8974942b-412e-4800-887f-6fa99e2acc01 · inbound

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars cites this paper.

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.396461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:38:28.718164Z digest=sha256:758c585f4079cc37c2a58549d81b06b15f90f640525e96b3e588a63aeab50ff4

Observation bee2084b-cdfb-4f8b-9679-bb172797f891 · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:17.845680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:17.845680Z digest=sha256:3f4b91ab3214abfe1eed29201dea3c928384acbb71b85608027dc3c420a582ec