Pith. sign in

Paper Citation Record · LEDGER

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2507.20091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20091 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:19.340049Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:49:26.507281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:46:15.325325Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dee9aa13-00b1-4644-b147-bbf0eba5df4f · outbound

This paper cites write newline.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.095238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.095238Z digest=sha256:3d5ef37388f29f095aa30556b6a60fbde711ee76d04ce2563d679dfa85a0af95

Observation 7d4af0ee-20f8-4091-a709-a63d1d3f8468 · outbound

This paper cites Dm-codec: Distilling multimodal representations for speech tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Dm-codec: Distilling multimodal representations for speech tokenization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.101576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.101576Z digest=sha256:adbab59321f370fdab2bd32bea76026cd33ce158c1795fedc8756a7eb6d9f2ac

Observation c92597f9-8e37-453a-b3f5-2f29d8379e52 · outbound

This paper cites dMel: Speech Tokenization made Simple.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models dMel: Speech Tokenization made Simple

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.106677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.106677Z digest=sha256:adafe6879c8ba2b9647fa5110ca792290fb2d585d113750d06d8a2f8f3fd66b7

Observation dddf75bb-7c71-40a6-9544-69c2e8c066c2 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Audiolm: a language modeling approach to audio generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.112745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.112745Z digest=sha256:0468944e9d700cfaca3fe08c71058267cc1ff77511036a181cf69419ccdb7893

Observation 6d818c27-e966-48ff-85f2-dccce0497ae0 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SoundStorm: Efficient Parallel Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.118430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.118430Z digest=sha256:c3f7f97addf08f08afe02b86da5e9a40a200eb1343a944fe637ee8f6d68031ae

Observation 0e085d97-ffc2-4cef-b51c-8f5fe7d5dae0 · outbound

This paper cites Giveness, contrasitiveness, definiteness, subjects, topics, and point of view.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Giveness, contrasitiveness, definiteness, subjects, topics, and point of view

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.212182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.123781Z digest=sha256:c8aafbbf1337f504a844be02c8797f1a6c6323b5e3ba76d4747df8a5b243cdc1

Observation 3d1e0e8c-5558-49b9-85ea-5876e630df6e · outbound

This paper cites DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.129079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.129079Z digest=sha256:7b813daf2383b24e7c739be6634676c19db007591b392d663d9643b4038b13a4

Observation ccb609e8-ddf5-4482-843d-273d1b8e377a · outbound

This paper cites DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.134392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.134392Z digest=sha256:bbf082c69b49cec0caf4189c4eb99d2af26259192d46d26bddd4904a29a06d7b

Observation 24823737-1366-4758-af51-1ecf7adb7e91 · outbound

This paper cites EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.140208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.140208Z digest=sha256:f579637d5636d8882d450cdcc1c9cfe81218afce31518b56cd060761ff46c51d

Observation 713a6d3e-8e89-452f-9f2a-83960024e07b · outbound

This paper cites High Fidelity Neural Audio Compression.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.145484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.145484Z digest=sha256:21c9be4c2359fd2fd969384f2065b28c3c787bc0ee3b14a83c7aa9761b3f8afd

Observation 7fccdd45-b965-454c-8f5b-97c77fc1e9f2 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.150827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.150827Z digest=sha256:40bf72e407601f9cd458cfd054867c1aa28f00dc793584583453836148029c28

Observation e274388a-c2c9-48b9-8754-7b5b1eedbe3c · outbound

This paper cites Elevenlabs voice generation platform.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Elevenlabs voice generation platform

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.195688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.156119Z digest=sha256:c0b7250dd39156f4612ff06d72e0e09cd1dfa2269242faa261c74a676a187c99

Observation 5612812a-b60a-4f0a-9d37-dae60a462b24 · outbound

This paper cites Recent advances in discrete speech tokens: A review.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Recent advances in discrete speech tokens: A review

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.161618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.161618Z digest=sha256:f6d2d18cd70df5d650844556481a0ec5ba6c518cde6cc51d835cdcac10da3506

Observation b0d96cdc-cac7-4247-9e64-4290d0ad8f7d · outbound

This paper cites Lora: Low-rank adaptation of large language models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Lora: Low-rank adaptation of large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.166411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.166411Z digest=sha256:decdb70bf1ca93b18119c4da734cfed9ca522192f1138bfb81c723ceb2ace475

Observation 285c13b6-7cfc-4155-a2fc-72d0f0da6136 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.171625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.171625Z digest=sha256:66348545048f8155cbca7e606ccc28f8c3bbb147d1c9657df76a88c120ed6d74

Observation b2f588c9-eb8b-47b3-8819-c41769fd9cdc · outbound

This paper cites RepCodec: A Speech Representation Codec for Speech Tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models RepCodec: A Speech Representation Codec for Speech Tokenization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.176618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.176618Z digest=sha256:c173cc79bc1adb4f8cbd1b6f3337d6551461dde820ea52b2c6b8635d8784c4e3

Observation 73502204-a433-4ace-8bee-40ce7176eac5 · outbound

This paper cites Crossing the uncanny valley of conversational voice, 2025.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Crossing the uncanny valley of conversational voice, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.170020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.181649Z digest=sha256:57aea0aefee3f55c20b451bb3525b9bb23ddd2322b0aef030248a8b740f5c3a9

Observation 357c6328-21d4-44b9-a094-d50b74a4d917 · outbound

This paper cites An open source emotional speech corpus for human robot interaction applications.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An open source emotional speech corpus for human robot interaction applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.186416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.186416Z digest=sha256:7ba439e7c493ac3208c37f0dd62a4212ea043e06495ba57a12cace2ba0f583a3

Observation e25b62d3-fec4-4b4c-aa00-4d0ed08db09c · outbound

This paper cites Style Mixture of Experts for Expressive Text-To-Speech Synthesis.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Style Mixture of Experts for Expressive Text-To-Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.191217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.191217Z digest=sha256:2a22e39028d70e72a59dfbc08777e3259e003eb4773dd5e3decd1d5a0993f543

Observation 4c2ec780-f37a-4266-a023-2e9471ff796d · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.196492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.196492Z digest=sha256:eeb5c096d20c71badddfcdfa6dc8707724d1dc7955ef5e5a303f287bf631d11e

Observation b7da3462-7686-40ce-81f2-9586298d6314 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Libri-light: A benchmark for asr with limited or no supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.201637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.201637Z digest=sha256:0eeb9b7fb9a940186bd4ffc119a65e9366be576f89fa4f01d9d337f32052f0d1

Observation 8995600a-5152-4815-b6a7-4fd51b794054 · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.206278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.206278Z digest=sha256:3f00d064cbafb490b213c431ca173f775db9d5b3d0010ce510bd3241c6b5397d

Observation d243f73c-5256-410b-8999-06111fd74af7 · outbound

This paper cites On generative spoken language modeling from raw audio.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models On generative spoken language modeling from raw audio

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.211165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.211165Z digest=sha256:530531e2b7fae592acb5c37706b63ed50d1b5914cdc82cd7f11f3956047ed417

Observation b6e6aa01-e127-43e4-b570-f2ed7cc28c2f · outbound

This paper cites Whisma: A speech-llm to perform zero-shot spoken language understanding.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Whisma: A speech-llm to perform zero-shot spoken language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.125620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.215741Z digest=sha256:7db3eb7a16f2b6991844a1ec2585371b1fdfbb526bc2b1b53f1b32aebf1f3005

Observation f5543cb5-b5ab-4813-95be-6a0a387dfe0b · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.109798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.220379Z digest=sha256:ae03cec1b2188c2b8deea44d147827227c8840afc8742f8c17a3be8592722144

Observation dae726de-a84f-464e-a818-55039915a068 · outbound

This paper cites Generative spoken dialogue language modeling.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Generative spoken dialogue language modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.094079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.225998Z digest=sha256:ae8039db7179f5d52f115b9740bbf99809a9641c11a78d7db6e634976f9d19b3

Observation f05a6762-2cf0-4e7b-a310-5d3613d603ee · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Spirit-lm: Interleaved spoken and written language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.078237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.230893Z digest=sha256:60e87004bc0904861c59ab97bbfce989ad37e3ec5e09f957bd55389e7d1f64a9

Observation 1da4e900-f7a8-4349-9c57-8b093de1d1f8 · outbound

This paper cites Long-Form Speech Generation with Spoken Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Long-Form Speech Generation with Spoken Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.236403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.236403Z digest=sha256:d72639dfe904c160af03c694b91804d4dfa7d873ea560ca89f1c2de0ba48ad1a

Observation d9db9bdc-d2cf-4d6a-bbd2-6002769aedf8 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Robust speech recognition via large-scale weak supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.241198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.241198Z digest=sha256:34a3960393a8a16f15fec207b557f093133f3933eb6f3412f38983b9bbbd8955

Observation 912c46e5-47a8-4262-a262-cbcd8c6d09ab · outbound

This paper cites Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.052853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.246565Z digest=sha256:8500a79361820996d2faf7717115f82b0edbb59bd7cd94ade16bcce06ed7df7f

Observation 8c198c27-1d04-46e8-90d3-e93b341966b3 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.251401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.251401Z digest=sha256:70a50a8624983eb7dc513100871fe5ba3add6f2989117e4991543cb54093d7ab

Observation 57f1eb8c-049b-4b88-b50c-66b3bf03ffe1 · outbound

This paper cites Shechtman, S.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Shechtman, S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.036673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.256580Z digest=sha256:6d7bf02f424189021d6e73f45bcc1b0970fbf187010b58e5bbfd9887e75f57bf

Observation 6fd81e2d-7668-4d24-bc95-5e4ec3d86d5b · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.261079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.261079Z digest=sha256:29e5e2c6059b2f5a1a707f5447c8700c529e1a41b407fef88c6884639340c610

Observation 5d103bf7-035d-42bb-9beb-54c4f55b5afa · outbound

This paper cites An analysis of the use of qualifications on the amazon mechanical turk online labor market.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An analysis of the use of qualifications on the amazon mechanical turk online labor market

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:54:20.020204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:54:19.266649Z digest=sha256:f5379ba3994622724491eba5f0dc4584775f698ee4446d9804d24502d9bb7d2f

Observation 5e32e358-db83-4460-8e6b-db9fff747a62 · outbound

This paper cites LAST: Language Model Aware Speech Tokenization.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models LAST: Language Model Aware Speech Tokenization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.271586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.271586Z digest=sha256:0323b5772256f9986fb23260fd90c9741fb12e625f22fc9927be621692d8bd23

Observation b94787c0-2378-45a2-b9a6-351247b3a200 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.277804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.277804Z digest=sha256:c554ace69f3dd4b252ab947c82a6eed0a45b4b7892b8881a9f37b23ecead5de7

Observation c4e6ed73-6450-4012-9043-9e684426521a · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.282895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.282895Z digest=sha256:a86f18a736180aa378fce2d25afa5dacbbb04749600948a7d8fe035ccd090827

Observation 20da324c-57e1-408d-81d9-26b3b92de506 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.288620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.288620Z digest=sha256:340fece8b5b86a2d1096f13a6c683584830a1d25ac6a424bd443c893cb4ab629

Observation 0ccb4176-9e30-49ea-9647-bda5a5e460e9 · outbound

This paper cites CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.293507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.293507Z digest=sha256:724fbcedec305183426064257beccf117003988cef6a55351da5e0f0badff188

Observation f66ab1d4-1d32-4423-afed-fe1d397bd964 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Soundstream: An end-to-end neural audio codec

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.298375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.298375Z digest=sha256:90180abfd24a20254ca026373c0fbbb6cebe4512577664aa46f8bd54a2764faa

Observation fe38e239-0e49-4449-b1f2-25255f4a6132 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.303523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.303523Z digest=sha256:31c010bc0933795097ea0b08bcb9935cf39a1ba6328f53999794e22c9f56e824

Observation f1f7c4c5-f1e1-4549-af60-efd7686a3e4c · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.309151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.309151Z digest=sha256:9cd0257c341327255abe1ed329186a9ab5c27892f1fa8e6fd6b078d642fdfe11

Observation 5b3a6cc5-aa00-47a4-b7af-0fbc3876389c · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.319277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.319277Z digest=sha256:f480c1ed302c737331a273657dad8e43bab42af839ae92d0df5cf32e6c73c277

Observation be3cf178-6cc1-4782-80f5-95f27d530a18 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.324529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.324529Z digest=sha256:df84135586685bddc1910a961bb5fa634f4855bbf33f4e282e1b75b0ef6c0a3b

Observation 5429ead6-1b82-4c58-b02b-1dea3c8f39f5 · outbound

This paper cites @esa (Ref.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.330235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.330235Z digest=sha256:996c04d0a67652839c2ff6b6c435310e089aead1a4cb8205367b7e002ab811c0

Observation 1c77e3bb-9bf3-4d61-80c0-397382e4feb9 · outbound

This paper cites an unresolved cited work.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.335195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.335195Z digest=sha256:0c2fb9666c62d2c29b1744d5bd5b3456fe1811216b788c18bb51a8c188e61452

Observation db04fd29-6e89-4cdd-9c95-6c6bac061f78 · outbound

This paper cites One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody.

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:19.340049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:19.340049Z digest=sha256:58065fbf9f6facef81cd0ebb1aca058e2b6360eeba44f875160c7d9def7bcc0d

Pith citing papers

Observation f56f709b-b828-435b-b0ee-482180d02993 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.327083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:a62a9897b24ebb22ce04804c6c63c03d0e3302cdc59a67666db829b40dcd14b5

Observation 73a32923-2621-4c7c-9c73-e2c763ebfd53 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.643075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:0edafb320ef93e871c8cbe73e3cf7d024c4db4de680626c9b49f44db1c387f2a