Pith. sign in

Paper Citation Record · LEDGER

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

As of 14 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 5 inbound Pith citation observations for arXiv:2501.01384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01384 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:32:05.104984Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T23:17:03.456746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:17:28.830734Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f033cc8-adac-4012-aae0-4cf27578d44c · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.897175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.897175Z digest=sha256:4fd289b68c078deae941f7716c94438eb3d62bd391e33d1a0fd3b7233320cbcb

Observation 0d6982a2-999e-4f38-8807-0b918d657330 · outbound

This paper cites Qwen2-Audio Technical Report.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen2-Audio Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.927050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.927050Z digest=sha256:c7f910d8bcb20a58c0af4bd91767046c4260b8b048b8bf520d99c726c75212c1

Observation ac489d20-9269-4d47-9b1f-015ee9aef61b · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.952630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.952630Z digest=sha256:4868bd9fdb9bfbc9e331442ba55cdb5f68acca739be9b07f7187c99da0bce001

Observation eea48f33-c7b0-4afa-98f3-7c95a493b137 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Dailytalk: Spoken dialogue dataset for conversational text-to-speech

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.788900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:04.973121Z digest=sha256:fc69734bd9c3bc6d48f49590d1abf6941323f58923cb6737f1c553add31e9d50

Observation a1bda65c-69da-42a3-9bee-ba9a7428ecda · outbound

This paper cites Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.980466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.980466Z digest=sha256:4d90919a43d38dcf363a763c4e9578b61d7f28bce9cb8c4e5970cbe3c39da328

Observation cab6075e-e96e-4126-8053-974c7a93de7b · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.987198Z digest=sha256:5ecb3f080756e9eb1f5c29b0a59c4fc94f9abb66d49608fc52fae0f1b5b67c11

Observation ee9a3be0-a176-45b5-bda9-b2bdf0e8327e · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.997721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.997721Z digest=sha256:81838522f4060efb72cfa5cf7293af23d35c29ac9a7e7ae506e0d18650b2970b

Observation afe84b7f-9da2-47d1-bc17-7dabddfe33cb · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.020670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.020670Z digest=sha256:106287c0281fbf387734fadd993a01e0d848bcb9d29376e921ff1289c4c590b5

Observation 9671ed34-5a79-4e7e-bbbb-c4cda352f985 · outbound

This paper cites A neural network approach to context- sensitive generation of conversational responses.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios A neural network approach to context- sensitive generation of conversational responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.692164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:05.039082Z digest=sha256:76902a3c40e4d388350a32dcab339a548212b30f45401451f4b851f049354471

Observation 1ad2cbf5-6b74-41b7-9b27-00e041e858cf · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.053591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.053591Z digest=sha256:aa4e90bed8a956c4d5bdf51503899779700f29bf490e0da102cdc92822ef04ad

Observation 88a7abdf-79f4-43e8-81b2-56563feb7325 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.059346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.059346Z digest=sha256:df86f93d4f998383db51ac1ce1c5fbb9734a770cc8d12b3610e76b7f6153dd23

Observation a5841562-877b-45c3-bd93-eb0892b3e020 · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.068528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.068528Z digest=sha256:8b8f2017c60364dff12df5b2a5c80cdb453a093e4c8adcfb4f8769fbc044e3af

Observation e43f5829-5526-46f4-9607-2323038d955b · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios AIR-bench: Benchmarking large audio-language models via generative comprehension

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.661560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:05.076706Z digest=sha256:dda32ce9296f7a39a6f7d8228580ba52b4db32f66f72d3cc3d1a601e8c0dc291

Observation adc1fbc1-a86d-44fb-aee8-b0fc8579a74b · outbound

This paper cites URL https://aclanthology.org/2024.acl-long.109.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios URL https://aclanthology.org/2024.acl-long.109

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.634168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:05.086658Z digest=sha256:931eee6cc91c9a3d332a4530309fe83197a7c80a4e7a5edd899e26b2bc2b678b

Observation 01ab8409-5a99-4ad8-afe2-e538f4505f41 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.104984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.104984Z digest=sha256:dde301a3f8b5b3a3af1058c257eaede7ebf190fd50a579c96ba01daf2d6b8958

Observation 877814fe-560f-402c-b46d-27869174b329 · outbound

This paper cites The cocktail fork problem: Three-stem audio separation for real-world soundtracks.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The cocktail fork problem: Three-stem audio separation for real-world soundtracks

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.769732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:05.003775Z digest=sha256:153af0a8c9c5941b1cb7f43ed25f06cfb0d7fa3be59bbd1cd706b892333eaab0

Observation 5dc1f954-88bb-4038-835a-2cee7b608670 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.935843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.935843Z digest=sha256:0b0899a7e52cf179bb913ec6b80199115f068ad38069370ee740a3edd6a47ce4

Observation 52a13339-d729-401b-a598-487c74ecd778 · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.906306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.906306Z digest=sha256:60f7c4165282743f6e3d5a08ed2d4b8097d691bfdafa5c154076749f88110f71

Observation 6c7903c7-8257-44d3-9781-757f4e9e1fec · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audiocaps: Generating captions for audios in the wild

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.808751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:04.958676Z digest=sha256:07776f31a8579816e564273ddc3a48d578dbf446d9263c9664ece7dd52e71406

Observation d88343bf-aec6-4c1d-8fe9-89664e868fb1 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.045777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.045777Z digest=sha256:b5fd725849ad10fd90bfc3f01e480d9c019ffc3e12c12ee5ed3f6953adea373a

Observation f3b5d209-1f93-46e1-82d0-411060b47e74 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.095673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.095673Z digest=sha256:a660d433d0d0b030b52717855bea4535da79b6dd404cfae524d5416b50744ef7

Observation fbdf322f-519a-4dba-b972-20fc0eb0faa9 · outbound

This paper cites End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.713573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:05.028644Z digest=sha256:332cbeb348e40aa35d23f498fe5343a98391541e9ea5b88b5f819f215ab61ae5

Observation b9183a36-597e-44dd-b039-53690bfd8af4 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.966190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.966190Z digest=sha256:4dd04efec1376106f649a6aa135017ff53f976187f2dbe6f2255e0face1bd2d5

Observation 80b7864c-4b9e-420f-8ed2-081b2ee94f50 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diariza- tion.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Powerset multi-class cross entropy loss for neural speaker diariza- tion

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.744280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:32:05.011324Z digest=sha256:1e0dca181e5513008afa2533736422b2df4a3ab67ce29890cc41dddfadf9c4ac

Observation 988a7467-632c-4296-85ca-059b6b981f90 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.916461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.916461Z digest=sha256:ee077c2d86a3f753f093e16b9a6783a2188e924955e8ac35e57d4f06cf32eeba

Observation 11f15a96-f792-494b-9a28-65498dc40a30 · outbound

This paper cites The Llama 3 Herd of Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.944578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.944578Z digest=sha256:843e432fe7b194a08d501868a373f3a867051aff6ff1e60e1e5f55612943657f

Pith citing papers

Observation 98d0d316-9867-42b6-906e-0f4a5197af4f · inbound

DiscussLLM: Teaching Large Language Models When to Speak cites this paper.

DiscussLLM: Teaching Large Language Models When to Speak OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.302323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T22:09:03.740109Z digest=sha256:0ef41cfb0f98907a441d307cb0d24374791c506a99d1d51941dda659b66dd6c7

Observation 591a55c2-6b29-4495-81f7-6d9c3820f334 · inbound

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs cites this paper.

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.501676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:52:43.396119Z digest=sha256:277510c8c6021060c535a063f0c6981d2bd52ed33c8701a7dcfcb9029b974f30

Observation 34137ee2-642b-454d-b847-092e5d8a2f04 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:17.001064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:ab89b85f387c3cd137e33212630cfaf48b86b4699d7d821a4d6d757fe3113018

Observation 166861d5-7603-435a-95a3-54759ec464b2 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.832711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:b3cdb1e8816acdd8e7636513f21d21416790986376f090cdd2ba95ddc9b38959

Observation 719381de-8387-42fb-8ead-4d789199210c · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:16:53.731741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:43feafb51931555a39b005821332000ae1befb5f456451b4e781cb5236f8514c