Pith. sign in

Paper Citation Record · LEDGER

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2507.20091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20091 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:19.340049Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:49:26.507281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:46:15.325325Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dee9aa13-00b1-4644-b147-bbf0eba5df4f · outbound

This paper cites write newline.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.095238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.095238Z digest=sha256:a3b30a7a5903806dfa621b5bb1b106ea583ce882e2d78a7b1f2885193df396b6

Observation 7d4af0ee-20f8-4091-a709-a63d1d3f8468 · outbound

This paper cites Dm-codec: Distilling multimodal representations for speech tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Dm-codec: Distilling multimodal representations for speech tokenization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.101576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.101576Z digest=sha256:73a0c7fb8fe6cf7f114adabe420295b565f1a945899d52c30577bf9080418094

Observation c92597f9-8e37-453a-b3f5-2f29d8379e52 · outbound

This paper cites dMel: Speech Tokenization made Simple.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models dMel: Speech Tokenization made Simple

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.106677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.106677Z digest=sha256:d23b635a38c79df7f57e36ac86b24e5dfc3107d19fd8e2de65a797ee18f57cac

Observation dddf75bb-7c71-40a6-9544-69c2e8c066c2 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Audiolm: a language modeling approach to audio generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.112745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.112745Z digest=sha256:d25d0c7a82783ec8bc2007f5d8c1780eec7c11155837881583baa755b75f56d8

Observation 6d818c27-e966-48ff-85f2-dccce0497ae0 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SoundStorm: Efficient Parallel Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.118430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.118430Z digest=sha256:72d0b9dc6d757cb4333533f4f2d323be279d39f86734baeaa2bf5353174bfdce

Observation 0e085d97-ffc2-4cef-b51c-8f5fe7d5dae0 · outbound

This paper cites Giveness, contrasitiveness, definiteness, subjects, topics, and point of view.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Giveness, contrasitiveness, definiteness, subjects, topics, and point of view

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.212182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.123781Z digest=sha256:cd7687da16bce03ffe5bcca33cae00375679a0f7d287044f9f9a5f86cb3420a0

Observation 3d1e0e8c-5558-49b9-85ea-5876e630df6e · outbound

This paper cites DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.129079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.129079Z digest=sha256:2a7166d67b6b84a26fc966a887b1829d215139f473b0322dc3f11dec0cc90c09

Observation ccb609e8-ddf5-4482-843d-273d1b8e377a · outbound

This paper cites DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.134392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.134392Z digest=sha256:55d819f11da78a451a681293d9cdee1671512bffa81c20c3e6b1a2a155001860

Observation 24823737-1366-4758-af51-1ecf7adb7e91 · outbound

This paper cites EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.140208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.140208Z digest=sha256:d200578eb2fb0f240fbf5fc06644868e35c7ee92d2f12b794d402979f8682d5a

Observation 713a6d3e-8e89-452f-9f2a-83960024e07b · outbound

This paper cites High Fidelity Neural Audio Compression.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.145484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.145484Z digest=sha256:7e1686daf57c25ab96fe07fe18861dbadd98a3a259df96bd241f0fef9faec542

Observation 7fccdd45-b965-454c-8f5b-97c77fc1e9f2 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.150827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.150827Z digest=sha256:54f3cb3bb69e89d12c31f384f569a9e96998511466618b92c15223c94e85d3ea

Observation e274388a-c2c9-48b9-8754-7b5b1eedbe3c · outbound

This paper cites Elevenlabs voice generation platform.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Elevenlabs voice generation platform

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.195688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.156119Z digest=sha256:efd39788ed5a9b5729b535b07054c40ab43465882c563e7713579742c86ecfe5

Observation 5612812a-b60a-4f0a-9d37-dae60a462b24 · outbound

This paper cites Recent advances in discrete speech tokens: A review.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Recent advances in discrete speech tokens: A review

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.161618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.161618Z digest=sha256:869a09f6400761b909c5970accfa646907a221d4109547d8aff0be6db9a7b1aa

Observation b0d96cdc-cac7-4247-9e64-4290d0ad8f7d · outbound

This paper cites Lora: Low-rank adaptation of large language models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Lora: Low-rank adaptation of large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.166411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.166411Z digest=sha256:0f9bf666862f6bbe816d23b709f90358a0fa98d6bc50e06c72ef63dbca771888

Observation 285c13b6-7cfc-4155-a2fc-72d0f0da6136 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.171625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.171625Z digest=sha256:7b429ed742fb5e6e66a4b1369885bee3b9aac70d6b55705b9d18f7022583bf84

Observation b2f588c9-eb8b-47b3-8819-c41769fd9cdc · outbound

This paper cites RepCodec: A Speech Representation Codec for Speech Tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models RepCodec: A Speech Representation Codec for Speech Tokenization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.176618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.176618Z digest=sha256:b79fd63b20a2ba8d4aa1f9801207ce3d5825b89c3997dba93dc8f9c6f843f1da

Observation 73502204-a433-4ace-8bee-40ce7176eac5 · outbound

This paper cites Crossing the uncanny valley of conversational voice, 2025.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Crossing the uncanny valley of conversational voice, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.170020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.181649Z digest=sha256:9d06971c199ae99b1957af2b981dad33847ed7bd099774087288e9f5c70d15c6

Observation 357c6328-21d4-44b9-a094-d50b74a4d917 · outbound

This paper cites An open source emotional speech corpus for human robot interaction applications.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An open source emotional speech corpus for human robot interaction applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.186416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.186416Z digest=sha256:3f81c68a781ce3b4299a54527722b07ace36c5603e77c5691e4331617dc27cb8

Observation e25b62d3-fec4-4b4c-aa00-4d0ed08db09c · outbound

This paper cites Style Mixture of Experts for Expressive Text-To-Speech Synthesis.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Style Mixture of Experts for Expressive Text-To-Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.191217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.191217Z digest=sha256:53ca959e00246e14adffb38d9ab10c009cabc5b877efae14e2590205c7dcbb89

Observation 4c2ec780-f37a-4266-a023-2e9471ff796d · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.196492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.196492Z digest=sha256:d83c0d29f954079f2404518ea5417e8dec3c7098c00823624f63434a11248321

Observation b7da3462-7686-40ce-81f2-9586298d6314 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Libri-light: A benchmark for asr with limited or no supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.201637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.201637Z digest=sha256:0167d2abe39dd0ee5685bb7e77f777157e6e41c216f5655d9d353cc2dae968e5

Observation 8995600a-5152-4815-b6a7-4fd51b794054 · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.206278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.206278Z digest=sha256:d5ef117e0584e6c31c25e068c84f029e247f2638b3b564f04ab8be79170872a5

Observation d243f73c-5256-410b-8999-06111fd74af7 · outbound

This paper cites On generative spoken language modeling from raw audio.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models On generative spoken language modeling from raw audio

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.211165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.211165Z digest=sha256:4958ec3d872c5229701c693658b75f283c9d1d52d4355cb77665be8e70dd17d9

Observation b6e6aa01-e127-43e4-b570-f2ed7cc28c2f · outbound

This paper cites Whisma: A speech-llm to perform zero-shot spoken language understanding.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Whisma: A speech-llm to perform zero-shot spoken language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.125620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.215741Z digest=sha256:0b4f16034bbac5185deb19ebda10dabb3307d94d8a939f56d7fb87d4f8c5c108

Observation f5543cb5-b5ab-4813-95be-6a0a387dfe0b · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.109798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.220379Z digest=sha256:a1ce5fabaee0a1ae625b33934fe7ba69afd69e782430e0cc79f63310ba5d192e

Observation dae726de-a84f-464e-a818-55039915a068 · outbound

This paper cites Generative spoken dialogue language modeling.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Generative spoken dialogue language modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.094079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.225998Z digest=sha256:58b37a4ce671b3bcdbcffe4610b38abf44157679d5e5e6ed48a54e7a5d008bee

Observation f05a6762-2cf0-4e7b-a310-5d3613d603ee · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Spirit-lm: Interleaved spoken and written language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.078237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.230893Z digest=sha256:4e3506db42af0d0757040812110df9d737466ea29abdb3faa38aa8b63b6c28e5

Observation 1da4e900-f7a8-4349-9c57-8b093de1d1f8 · outbound

This paper cites Long-Form Speech Generation with Spoken Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Long-Form Speech Generation with Spoken Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.236403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.236403Z digest=sha256:f779f1e6657c9224f08568feb9d00bdec41fa6172b6d2764eb64974150c5eba4

Observation d9db9bdc-d2cf-4d6a-bbd2-6002769aedf8 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Robust speech recognition via large-scale weak supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.241198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.241198Z digest=sha256:45ff05a99eb83c3ecd07d62de7b0cb132263807bf90fdb14ab61de88a54fe7a1

Observation 912c46e5-47a8-4262-a262-cbcd8c6d09ab · outbound

This paper cites Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.052853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.246565Z digest=sha256:03b4c17452547a8d60cc0a9f5e61000c0c6489a3d62d5b642c8c8edd773fd036

Observation 8c198c27-1d04-46e8-90d3-e93b341966b3 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.251401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.251401Z digest=sha256:21e8d56751d4b56cbdc05661024d9d7b0bccf0c06f413045c8f806553ac8054c

Observation 57f1eb8c-049b-4b88-b50c-66b3bf03ffe1 · outbound

This paper cites Shechtman, S.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Shechtman, S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.036673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.256580Z digest=sha256:dd5ee0d62087fe31cfd3fcbe47a50ec96e5f3ba4f603809478a042bdb188d999

Observation 6fd81e2d-7668-4d24-bc95-5e4ec3d86d5b · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.261079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.261079Z digest=sha256:ec3cf38d5280afebc9698086281bba1cb6215d628fe944ab58cbbc7a0b32e2f5

Observation 5d103bf7-035d-42bb-9beb-54c4f55b5afa · outbound

This paper cites An analysis of the use of qualifications on the amazon mechanical turk online labor market.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An analysis of the use of qualifications on the amazon mechanical turk online labor market

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.020204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.266649Z digest=sha256:1920cdc5364f48b2f6a75b916a9b1c7141db66bf59fc3fe92540baa0c69cdb7c

Observation 5e32e358-db83-4460-8e6b-db9fff747a62 · outbound

This paper cites LAST: Language Model Aware Speech Tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models LAST: Language Model Aware Speech Tokenization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.271586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.271586Z digest=sha256:4702255e18ed3775166942c72459830ca1383c0389fa5cb64272364ee0a3d1f5

Observation b94787c0-2378-45a2-b9a6-351247b3a200 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.277804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.277804Z digest=sha256:48e98a90c551b78503d36259f63a303f0489b3b5d0e6f6723f17399f3d168020

Observation c4e6ed73-6450-4012-9043-9e684426521a · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.282895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.282895Z digest=sha256:68e09f3f79eec7ec2a7328ca90a7eed3c4479ade3718710aab80c6b3304ba96d

Observation 20da324c-57e1-408d-81d9-26b3b92de506 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.288620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.288620Z digest=sha256:6085a28a5a6c6910b06a762de186e91a0093e33ea5de377af8202db6a64e89f8

Observation 0ccb4176-9e30-49ea-9647-bda5a5e460e9 · outbound

This paper cites CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.293507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.293507Z digest=sha256:3c1f1a21aab0c240e81e411d4642413f503fee2b9b67714b3d348e636829bf34

Observation f66ab1d4-1d32-4423-afed-fe1d397bd964 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Soundstream: An end-to-end neural audio codec

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.298375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.298375Z digest=sha256:20cefeb5686c8f2e9370f402895fae12119817c122831f824c4bb699ec72fbdd

Observation fe38e239-0e49-4449-b1f2-25255f4a6132 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.303523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.303523Z digest=sha256:54b767be2bb07f611af8f34c42a966856ec1139175145e21405f0b91520fa906

Observation f1f7c4c5-f1e1-4549-af60-efd7686a3e4c · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.309151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.309151Z digest=sha256:2e5b6fc42514fccdd030183deceab58a00bf4036fbcf033bbe4ae9ce20d55ad9

Observation 5b3a6cc5-aa00-47a4-b7af-0fbc3876389c · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.319277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.319277Z digest=sha256:30d4f8a408df91e38b4847410e3050535a2b5e6432236cc68a0e41cad8b25e4d

Observation be3cf178-6cc1-4782-80f5-95f27d530a18 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.324529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.324529Z digest=sha256:7a2fd77ccd5c36fc8cae514508214fc5ecdcc430fc4a095e765a3949ad896cea

Observation 5429ead6-1b82-4c58-b02b-1dea3c8f39f5 · outbound

This paper cites @esa (Ref.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.330235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.330235Z digest=sha256:2e7a3cf9599057416e1e835cf35e6350d8d30f71aab51ab2183c5a72fa1cef1c

Observation 1c77e3bb-9bf3-4d61-80c0-397382e4feb9 · outbound

This paper cites an unresolved cited work.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.335195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.335195Z digest=sha256:7a2f770504de4604b0af291e77ada6a9ad0b752284cbca885b03e34b97f02040

Observation db04fd29-6e89-4cdd-9c95-6c6bac061f78 · outbound

This paper cites One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.340049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.340049Z digest=sha256:2c2b0cdef7179ec1ba361f0771a0fbc3dd7b5beb20d1d4ba48a44ac959ffea99

Pith citing papers

Observation f56f709b-b828-435b-b0ee-482180d02993 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.327083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:21b475050713fbfa4a19d2467c1523f2fd1095223eb072ee97cb14b669df337d

Observation 73a32923-2621-4c7c-9c73-e2c763ebfd53 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.643075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:f95c037184f6deb1e9e4a84a141dfb06b8169f09cb9d89ff6a55a61c1d96b037