Pith. sign in

Paper Citation Record · LEDGER

Towards audio language modeling -- an overview

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.13236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13236 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:30:27.004124Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a7b5bfff-4835-4fca-9494-232bfbd952d1 · inbound

GenVC: Self-Supervised Zero-Shot Voice Conversion cites this paper.

GenVC: Self-Supervised Zero-Shot Voice Conversion Towards audio language modeling -- an overview

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.004124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.004124Z digest=sha256:b04efb9801f3481f409bdd0619772d46caaf479d2aa3570efa362aa18620b3be

Observation 152e2037-dd50-4766-900f-e45f6ed9c7d5 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Towards audio language modeling -- an overview

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.204219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:1c750f72fcb88651e8cfe1a03826e0cd99be6ece0d90fbb1a75f0f5601e892d9

Observation fb99c11b-297a-4068-8a44-b34c2e4f1c89 · inbound

Learning Sparsity for Effective and Efficient Music Performance Question Answering cites this paper.

Learning Sparsity for Effective and Efficient Music Performance Question Answering Towards audio language modeling -- an overview

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:45.820916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:45.820916Z digest=sha256:465c234262ef000d7730c4617867e59241fe2c7eed637fabdcf1c3ba68093497

Observation ec034de6-3d8c-4ca7-9e9f-86ae48eb45fe · inbound

Towards Generalized Source Tracing for Codec-Based Deepfake Speech cites this paper.

Towards Generalized Source Tracing for Codec-Based Deepfake Speech Towards audio language modeling -- an overview

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:08.337092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:08.337092Z digest=sha256:3b50c8544056c24682bfbc643b269aa311eebee31c9191e97591eb3494d689c9

Observation 429258a5-fdab-4c66-94c3-0078deb08521 · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Towards audio language modeling -- an overview

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.281078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.281078Z digest=sha256:5ef4ff5f7186588dbbd90fe46b75f35a0b4953766c59311d319c0a4815f77214

Observation 9d967d19-e2ff-4963-b199-fdb9b9f2e686 · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review Towards audio language modeling -- an overview

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.301818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.301818Z digest=sha256:e7046258dec704124732196e5107069e3c2904ce256e379d79bfbb0d0565edd8

Observation cd07aad6-6f2d-4474-86e0-6e737ebc5f75 · inbound

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine cites this paper.

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine Towards audio language modeling -- an overview

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:34.416786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:47:34.416786Z digest=sha256:7186365f1f99f33d69f48f77d29e02288feab26f9e7e7e25d9298c968c0a6b16

Observation 536b2ec0-62b9-481b-b8ee-ab4ccc1042b9 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Towards audio language modeling -- an overview

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.834220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.834220Z digest=sha256:cfa497ef0ea74ad5ce306ac0ce3a04f2d760a198df0f6e489e338adc6429300d

Observation b5f5294e-c570-4d40-a952-cbe42cef73c7 · inbound

Analysing the Language of Neural Audio Codecs cites this paper.

Analysing the Language of Neural Audio Codecs Towards audio language modeling -- an overview

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:38:50.683004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:38:50.683004Z digest=sha256:7715ea25538f99696423db8112c4205a29dc02ce2fe6a7ee0db59dc5edd3e2cf

Observation c061f0f4-697c-43f6-b711-8f732d38abb6 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos Towards audio language modeling -- an overview

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:37.073486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:37.073486Z digest=sha256:d7736df5cbb37b48632330b0aa9b444ea0138ce203e8bd707156cc0b0479c861

Observation 341cae12-85cc-494f-8067-3c19f14c93e3 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Towards audio language modeling -- an overview

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.326589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.326589Z digest=sha256:a01e6445acc76ce4415bffcb3917e569ed4456746203356866a4980ea6017cd6

Observation b7adefb2-5874-4e8b-8a3e-d4487ece71a4 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models Towards audio language modeling -- an overview

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:52:35.713570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:e884abadce7470151c09af61ebc501a9dc756079846d056c399c5996bdf066d1

Observation f348969f-392c-4c6a-b56a-49936e6b55ae · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation Towards audio language modeling -- an overview

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:02.483589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:db594ce0b8df4fb18087c111cee85a9f7b1eb3e0cf74b370e73a0ada5a947de1

Observation 2d628b51-8860-41ea-8eb3-39352ed9185c · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Towards audio language modeling -- an overview

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:41:26.269011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:9745f7ddc160712a688d6de548a992164b549b232edaaa5f6e23787d2b4922b0

Observation 1411c4a8-c068-4e55-89ae-5d38c5aeae59 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Towards audio language modeling -- an overview

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:16:08.854881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:51ef5b4a34edff4fe3ef6e7d9b56a71a4248081cab315526b888f700783f170d

Observation 149ddc70-4168-4d74-8ef2-2fb779029eb3 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Towards audio language modeling -- an overview

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:15:07.870811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:fdfb3be2b8ce7baef1429d4de70db353b00d8b53ab6b9e9fe219e408472a638c

Observation aa31f89c-0fb2-4faa-9a36-5c53ade57752 · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects Towards audio language modeling -- an overview

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.825005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:d6691a3a2802cd69fcd6a0fe30fb0ebf6011fce4be9ca0631e57ea8c29c0c2d0

Observation 31861423-8f53-4dba-9350-f132a457214d · inbound

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech cites this paper.

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech Towards audio language modeling -- an overview

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:17:21.989251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T20:43:29.213147Z digest=sha256:00dc6cad6d8349899bbdf7fd5d2156f806ceae257d574065d4aa1e361abe0925

Observation 922edaf7-1f02-4279-acfe-b2a0d97960cf · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Towards audio language modeling -- an overview

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.306809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:626a2de1b07df5df3c97671975f73f3e24cedb206b7bd3206636b995fb5d4062