Pith. sign in

Paper Citation Record · LEDGER

Sparks of Large Audio Models: A Survey and Outlook

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2308.12792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.12792 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:02.087260Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T20:45:08.153364Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d0ea79f4-c8d4-4ab4-843d-bd20c128ea11 · inbound

How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System? cites this paper.

How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System? Sparks of Large Audio Models: A Survey and Outlook

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:04.504086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:04.504086Z digest=sha256:d0905c25ee0b2cbb09c34301fedb25ccccedbbe38ba055a4c35ab886e48ff1df

Observation 3b48b721-800e-443b-ad35-198a71042cd6 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Sparks of Large Audio Models: A Survey and Outlook

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.087260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.087260Z digest=sha256:5efe3e6ab07bdfe4f6a7c24efea767dd1db9ac6346ed8d64f40a08ce3c55ff74

Observation 8799987d-4c64-45a9-be1c-410a45cd3d9c · inbound

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison cites this paper.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Sparks of Large Audio Models: A Survey and Outlook

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.658865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.658865Z digest=sha256:13abe2f1389a4b1d841795e74808d5171defbbb1ae81acd9b50a4b507e7a74f7

Observation c5586769-663d-4e7a-8a58-853b186df7fa · inbound

From Screens to Scenes: A Survey of Embodied AI in Healthcare cites this paper.

From Screens to Scenes: A Survey of Embodied AI in Healthcare Sparks of Large Audio Models: A Survey and Outlook

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-10T20:43:10.250689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:43:10.250689Z digest=sha256:6be615ca6bd782cad20b9e33138b30abf7e93433988ad6b7b1dae2c53e411c01

Observation 348d5f4f-9348-4e04-bbce-ceabc736624b · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Sparks of Large Audio Models: A Survey and Outlook

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.156543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:de014460e0c58ed54bb18db000220c021a107aca33847848c4d4c3e68667a8fb

Observation 25ba0e21-4489-4a94-aef1-40529f76f3ca · inbound

Probing the Robustness Properties of Neural Speech Codecs cites this paper.

Probing the Robustness Properties of Neural Speech Codecs Sparks of Large Audio Models: A Survey and Outlook

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:31:36.944186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:31:36.944186Z digest=sha256:44858a6ba7d69dfced26b5471332848eb882d9573620d88d6c3efc82fcd40c01

Observation b81e8bff-6915-42ea-a251-eb248a06c311 · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Sparks of Large Audio Models: A Survey and Outlook

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.148862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.148862Z digest=sha256:1ac066b2ee5e6a5cd5c2495a13b1fd406a15498b77fe9e247e08f58ed755b150

Observation e0d309b6-211e-4ad6-974a-af2093367958 · inbound

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents cites this paper.

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents Sparks of Large Audio Models: A Survey and Outlook

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:45.100086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:45.100086Z digest=sha256:7666b57c358501fd0a93d9fe0469c5031933b1679e86f5dd5bc64641c640e60a

Observation 52fed0d3-eb66-4725-b249-4535076df960 · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan Sparks of Large Audio Models: A Survey and Outlook

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.515683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.515683Z digest=sha256:0f1e4d4e18c0808e6ebc1cafe0d96c64cfd9d69c0cb1af40738cd8cd138d866b

Observation 1e2b93cb-934e-4adc-9cc4-32e48ebf169d · inbound

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation cites this paper.

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation Sparks of Large Audio Models: A Survey and Outlook

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:56.898532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:40:56.898532Z digest=sha256:1ea95a713a14b7f25cd092ed918f7e936942e209597e4bd5cbce65bc522a4f54

Observation 1e380dfc-05f5-4624-8b86-ba2db684e2f1 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models Sparks of Large Audio Models: A Survey and Outlook

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:52:35.681744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:31f7b1c5dca3c4c74a28e4e2640076845c4df5ed61cda90514cfbdd455c0abaa

Observation 1b5d64a3-e72e-4c99-85af-ec6faa9844d4 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Sparks of Large Audio Models: A Survey and Outlook

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.776305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:df63b56657dd6087b3428d573e5ca9adb40b792efd2c348b5ddd39a01fdb9210

Observation 4a84b826-a364-475a-8450-b2583259cb58 · inbound

Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach cites this paper.

Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach Sparks of Large Audio Models: A Survey and Outlook

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:47:37.375859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T08:44:02.494895Z digest=sha256:5edd9985e0216eb6e429dbc286a06c86e43c33081d9e26186598075927c2bf33

Observation 1c77c32c-4c42-4db6-93c0-c645629de473 · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training Sparks of Large Audio Models: A Survey and Outlook

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:18:07.311169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:b0fab10e9e34d44f4fe54a2863158fded7494bb4df694bda57893456caedeba8

Observation 6c6d05f5-03a7-48d3-b83c-8c7ed597307e · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Sparks of Large Audio Models: A Survey and Outlook

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:48.960571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0ec2d292e9d2262c56c33e643544eb51721e02140e2332938d33e87d8f8dcf13

Observation ca5800a4-c99a-4c78-94eb-dd7fa1bf053a · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models Sparks of Large Audio Models: A Survey and Outlook

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:09:24.461469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:bae6f9adace53eb8ba654ef4b396067667c69d14b900875d4b8a12a544a7d12f