Pith. sign in

Paper Citation Record · LEDGER

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

As of 24 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 4 inbound Pith citation observations for arXiv:2412.01078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01078 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:46:16.782429Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:52:16.074875Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:24:03.552101Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b052b07-b8ff-4569-8727-922ce7be9484 · outbound

This paper cites this" or.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data this" or

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:16.985455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.760093Z digest=sha256:32eb2390d82f3e2a10352de08ec113b1768a53f2264da2855e43364d2aec43f5

Observation 7a01eef5-f9ec-4282-b3cf-a5c3ccfe2db8 · outbound

This paper cites Qwen2-Audio Technical Report.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Qwen2-Audio Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.729612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.729612Z digest=sha256:7462513f05238898b161903b794f59f244f9ed0b9146d5983c49abf4f713329d

Observation 91cf2943-42c7-470c-a3dc-3a76cd0cb803 · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.953950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.769330Z digest=sha256:e16af96a06fd022a3f66a9ed5a7cb139fa0e5985869e80630ccc30e48e6b90a0

Observation 27835fd9-0449-4821-bda1-84734b888dcf · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.739240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.739240Z digest=sha256:ff3c4cdc77ea22735ce20e0410b744c252789ad2d4ad258cf1b4d5bb27e188f3

Observation 83205df4-17bf-443e-9a74-7d2871d0382f · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.744798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.744798Z digest=sha256:95bf3a444369cc9847300930afe50b610eaf436883df1d71f7d33c6ccaa86d22

Observation 5aaaaffc-32d2-4273-977e-4108622727d1 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.749983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.749983Z digest=sha256:a96d36a250a8818f14962eef9e860ac0f23c75b96eae8df80c63dba67901ed1a

Observation 5efdad44-6388-4046-acd6-3c62a542dd26 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.755242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.755242Z digest=sha256:262d002f276d987b0cb80fb218cbc81abbdb52abff24e35ca1b4b3d07f56ea57

Observation 36698443-574c-4354-b9f7-88641718841d · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.968375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.764723Z digest=sha256:cb9d9a9ab976c445aeab054914cedcacd6a1cdde79f52370cf001f8d6f1c574f

Observation 1fdf82f8-b84e-45c2-a7b1-230f50630767 · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.939317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.774097Z digest=sha256:c17a72070a180caea635c367f0b14a145e46b0ec72d4b5a6f85290d49bd90b99

Observation d0fc32ca-08c7-43ad-97a6-c069c226c1a3 · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.924599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.778156Z digest=sha256:4ad55c8238373131d23f06009d6e0ce35dbea148124eec73d44f8d773fd8d54a

Observation 56bdf33f-6dbe-4b09-91c6-31a3af7406b5 · outbound

This paper cites is_suitable_for_speech.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data is_suitable_for_speech

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:16.909892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.782429Z digest=sha256:0182639f69ffec6f34acfb4681c7852a0fb6092279bf64cef463df94d0ace9db

Observation 88578298-d84e-4d76-9086-eb5c3100786c · outbound

This paper cites In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), pages 6968–6972.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), pages 6968–6972

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:17.001536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:46:16.734607Z digest=sha256:722052dc98d7a4c73a232645b124cabae5af5bce5c98ae3563b9e555ee97dea5

Observation 4e5744f6-03e9-4d65-8dd9-c5a8f5045879 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.724250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.724250Z digest=sha256:be93d86e354d8e32abcf7876a2d78fe3a5634306083fae30520f524982ed4652

Pith citing papers

Observation 46fcac79-c811-47f7-8aa0-85c9f5e7ab93 · inbound

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue cites this paper.

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:52:16.074875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:52:16.074875Z digest=sha256:692499d17caa74c184fa387df5d55cd770d59552100f16e109a693f80228fe88

Observation 69e3da9f-8225-4c8d-bd59-84db56e65a27 · inbound

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation cites this paper.

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:01:07.569586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T17:00:59.421441Z digest=sha256:1f7bcf717a74b602f2616e2eb21f59f5dde694007094e7778c5dcd03435e872e

Observation f8ff0edb-4d33-4701-92c7-17ec5418c873 · inbound

Voice "Cloning" is Style Transfer cites this paper.

Voice "Cloning" is Style Transfer Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:02:47.145012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T21:01:24.307137Z digest=sha256:537854de7a4b5b0c55604df683b48c9a76d837cf5b98aedda29ba7d9dfb80724

Observation a54bac18-3bdd-4ef9-9df4-cdabaf8bb88b · inbound

Voice "Cloning" is Style Transfer cites this paper.

Voice "Cloning" is Style Transfer Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:24:03.553542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T08:20:29.918036Z digest=sha256:611a7a61680be82da64fdfd3f62caf031b3aa460621f0a6e4240948b16c26262