Pith. sign in

Paper Citation Record · LEDGER

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 8 inbound Pith citation observations for arXiv:2508.04195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04195 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:51:32.702267Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:40:38.072150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:25:02.434510Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7329c7c-d665-42d1-83e9-0297e80d850d · outbound

This paper cites , " * write output.state after.block = add.period write newline.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.419290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.419290Z digest=sha256:6a8fcd89b03c34191a675f8f2615bad82f58b62024a1b8f1f02531e0a60b795b

Observation 72272a94-9074-45d3-9444-56577ac6960d · outbound

This paper cites write newline.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.470752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.470752Z digest=sha256:bd631e038308bca97578be150dd262968a84334636f4eea45dcf212467038248

Observation 001b74c3-9d4b-4ac2-b2fd-de1a1b7f24b0 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.524524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.524524Z digest=sha256:500f33ba55c0e585bfe968488c0ba8944dbfe4239ec24e0520636773f1114c7b

Observation bc2d1ba8-8745-4564-bf8f-7f57b24c0aa9 · outbound

This paper cites Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:51:33.246231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:29.579149Z digest=sha256:3712d10dfcfc662001a25abd50f66e7d89206395f02742f70af567591c6e097e

Observation f7251430-0160-4158-b836-a7496021ba59 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.678062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.678062Z digest=sha256:f64de7f0d061ef1cfa0ede81dd64aa2b5b54bdac5ebc3bb1d3ef5895059dc80a

Observation c2d1e50e-b985-4320-bbcb-8f6a59ffe200 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.737750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.737750Z digest=sha256:e1b95cc923c8e71a48f0aab8d5be1bcfd5fb6c463c4d5ee79d1a920312d36b1e

Observation 9ca561e2-bdbc-49fa-8171-9c771a950e8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.827602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.827602Z digest=sha256:32378b20e3a03953b6513265a98ee6a17e7383d3f8c274163f098949f2df0cf9

Observation c3fbd9e7-19a0-47af-93b3-4c819a8ee3e5 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.556797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:29.891439Z digest=sha256:610e9a781a959ba06b06d1b43e6017fe9796a8fc1e26d9a9098e77fbb9c71f09

Observation ea8f7aa4-6cb6-4a43-a588-7b5dfc2e2fbf · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.957770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.957770Z digest=sha256:32196ed9a603bc7f46e5da6e5e9cf7be628f3d0fa22356e717c9e42e8cb765ff

Observation b679bf31-2d50-4682-94bf-78ffa0bfacc2 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.009070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.009070Z digest=sha256:12cc85754ac49812c5b001104baa73dbde91fa1dc17167869b0318446732bcc8

Observation 15c7e2f3-41c4-4e96-ba2e-5ba58fe90610 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.106597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.106597Z digest=sha256:5827c134ca487f24305873107c5ed8588555d49a7dac6e40f3347a41535ea060

Observation 89df5f50-df60-457b-bc39-d774e3dc31b9 · outbound

This paper cites F.; Ellis, D.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations F.; Ellis, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:36.375230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.179764Z digest=sha256:cc4a1c6d356d5a84f27f063dbfcb89bef3157ba102152dc2912816859d7400d6

Observation 331361af-2510-447b-bd02-e88df15243df · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.216252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.195673Z digest=sha256:add4568f7ae487beeb73b303eed0cdb34509d1b28ccf2da117c93c0eebf881db

Observation 0deb0d0e-f8ba-4bf9-bad0-278bc699f7b0 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.042699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.296507Z digest=sha256:77dc037b251e075debbbd5ddb59210a4ab2afa567e64709a5e9aa34fe0e2bce3

Observation c671d657-7cf7-46fc-a3ff-b7c02cdf8e9e · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.398007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.398007Z digest=sha256:3b5fc2c5061c537ed2580cebec22871862864eda6103440ba8c1d23dec108db0

Observation 4b1c2fe0-0dfc-4965-84d1-e33b4bc65832 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.882856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.527531Z digest=sha256:077575e58a4f14ee4eb430fd6ae100eb9bc5aff069195bcbbbf15563046b534e

Observation edc42f4b-f738-489b-bf02-57c465a9e932 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.658204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.658204Z digest=sha256:7c73459e9afed8f0b4617fdc937f176751ab366869accfa058b8d90042ea2ddc

Observation ee47a100-26eb-4ef5-ba87-838361fe724a · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.769766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.769766Z digest=sha256:c1513dfdad082981a581df9f9db392a7dbd4da16193837ff170eb9b39a889362

Observation 3b946c59-8ff3-4d57-96c4-14438a6cac74 · outbound

This paper cites S.; and Aslin, R.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations S.; and Aslin, R

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.727651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.924211Z digest=sha256:f7d0f8aa2bf9efbd8cb59b60b8fcd18442f3a79bd05af1d3a2981d2cc4b8857e

Observation 6f58bd19-c46a-484a-ba33-d99fb23b5c44 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.582626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.118101Z digest=sha256:34748b4d3f7aaab339274bb8bbe2f490e26fb1f584e7d452c491e072714ecd63

Observation c48991fa-e5de-4190-a353-fa8ff95e2e32 · outbound

This paper cites M.; and Nass, C.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations M.; and Nass, C

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.440905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.272438Z digest=sha256:8656e47422eaf5c505612dc417db064b78d168015ea0e12cabd8792907317f13

Observation 7d22d4cb-c8e1-4550-a9c7-783d4027daf7 · outbound

This paper cites E.; Rohde, H.; and Corley, M.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations E.; Rohde, H.; and Corley, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.276295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.335779Z digest=sha256:31c0a5b1d80460462bd56e16f0ac9b9aea4c6eb2a0e2d4a73b5c269e6703d487

Observation 4402f487-40f7-461f-be2d-02e40eaa173d · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.393989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.393989Z digest=sha256:13462b95b2165f8bda84a63f4b3bf6e055b9c702650e1d9c67c34b91e9b8c689

Observation 90a046bd-2328-42d7-be3b-53404aae77a4 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.131919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.467268Z digest=sha256:08cb25e22f0b72236ca4d2110a1ecd80c798e4b76e62ea78f9b7f4d2ed80874e

Observation 4d568a58-48f6-438a-bb20-7a83a32d2c25 · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.540103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.540103Z digest=sha256:035ac9e7de453af457580928bc89007841815d7d589cbbdb03c5ba947cbd991b

Observation e61ee274-39f3-4365-b648-63172564e471 · outbound

This paper cites M.; Li, G.; and Du, C.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations M.; Li, G.; and Du, C

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:34.968392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.638027Z digest=sha256:3edc280f6490a7a717038034b109caf1fa015b390944356ab72e6471d64d5b57

Observation f1c1672f-6204-4d98-aea6-49563bb65895 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.794508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.749455Z digest=sha256:0aed68a40b280363f151b02a972695eaeb0279d8540ece0b9018a82d9e880d49

Observation 2ac5f453-ff97-419e-b111-48e6799535e2 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.866776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.866776Z digest=sha256:9226b69d25500cacb9623440a7f6277461f9d99923ed8424d2604c319e97cb56

Observation 350cd4e5-6116-4810-85fa-692f7834aa6d · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.623200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.939646Z digest=sha256:b17cb97ae2b8f3a436142e130c9c38548d227336d3cdaa2c5bb4d51e8e549e2a

Observation 85795783-f8c7-4e9e-9ce3-5e4595ae444a · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.396959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.025999Z digest=sha256:06220acf3aee6330e59364e29fdae644fb35e644293904608fa4dd6af282e10f

Observation c679a6ea-d81e-483d-a059-f530be6fd47f · outbound

This paper cites A Survey on Neural Speech Synthesis.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations A Survey on Neural Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:32.131382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:32.131382Z digest=sha256:3cf813e8d011f0edd244da8c1e9f090fc5068a9a89fb5d3ca664c0df4f227e04

Observation c2447670-cbd6-4f8c-9fac-3720882b3d2a · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.094456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.233349Z digest=sha256:994f506482658c70e6eb27e0731c57760c23722451a8eea4e62fdedeac887ca7

Observation 24c5b1c2-1b12-4404-9cba-5c97d514214d · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.881745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.333796Z digest=sha256:1b9d56a8bf1973c85c13198c73d3b24ac33f438d2310297132437f57218328f6

Observation 701de64b-c30b-4506-bb16-5bd8000d593d · outbound

This paper cites DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:51:32.888098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.388095Z digest=sha256:1be4da775100656ac9696b61761c006009a470065250ef32c1f9728d79e57752

Observation fe9c6c71-7323-4996-bf15-f3cd272cfa16 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.665539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.480445Z digest=sha256:ce6d6353b2a0ac99c2be9374797b809e04d5ad69a69093dc21fc8115ee2e30ab

Observation e1eb9f8d-3177-42f9-a402-81b86bc9775d · outbound

This paper cites E.; Thakker, M.; Tompkins, D.; Tsai, C.-H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations E.; Thakker, M.; Tompkins, D.; Tsai, C.-H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:33.505371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.557490Z digest=sha256:5ca2e8c22cf60773656a4482b7b2819bb246eba465e9897bccda93be2a5922b0

Observation 352d1629-f1fb-4857-a0c5-0c74c2bd5a9d · outbound

This paper cites Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:32.620801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:32.620801Z digest=sha256:f92ffbaa0b3180cc0deb3172d710f401425c03ac71125691e301f0582d753bd9

Observation b3f16aa5-609e-407f-8418-baa6bc143a82 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.368415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.702267Z digest=sha256:3d1dc986cbfca50a0716fb853acd3b8a7bbb10d0da0384b1c899c574b559c265

Pith citing papers

Observation 4a89afb3-bc58-402c-a22f-e4a9c46b3b84 · inbound

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices cites this paper.

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:40:38.072150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:40:38.072150Z digest=sha256:fe2bd736a38bab58224d876e1d2db411162ad165f848e7b43048cd016c2a4b04

Observation dc465a96-f3af-4c05-bdc3-c67a12c398db · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.826968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:1e6ec4d22d97187ed0c948c02e7570d2e7e37944703730ebcdb434fd6decdec0

Observation 610ee6e7-2250-4543-8895-9e5295d4fd57 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:13.005685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:a4c006495d7b9c51199b162820f94d2b7bcd8fe44e01f93e061816e6486420c1

Observation ecc84b56-edce-420c-b44d-25ffcba4899c · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.976948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:4f17d139a24ec91c158a3ee918e8edee6bffea9c7f171ccf065ae31d39d29905

Observation 16435fef-12da-4d26-819a-82ef45fa8ac9 · inbound

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model cites this paper.

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:07:13.507801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T04:03:27.608638Z digest=sha256:cde12866d135ef5d09c11f539517beba6f1473f335276bd33750dbba49183767

Observation 8eb7ae9c-ae45-4db1-b2f8-41d6733a6577 · inbound

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR cites this paper.

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:37:29.604600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T00:29:39.365107Z digest=sha256:ab3cf1a0596c474327f39b316a753521c14db855db83a2229703d8b1713ea841

Observation 780fd3bc-2880-4093-a9f3-601039e19f43 · inbound

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios cites this paper.

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:25:02.436023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T14:20:24.828073Z digest=sha256:4c231e0302a35e24d9b8492c538bacf7671888f0e52250a01fe0f9edc6e72499

Observation d6960b66-d470-4ba4-9ba0-c4a9026b2694 · inbound

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing cites this paper.

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:01:52.696984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:01:52.696984Z digest=sha256:4a42d945ac4edfe4bb235ddd02f5c4c6876231533749e56cf2e61fd85acef58c