Pith. sign in

Paper Citation Record · LEDGER

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 7 inbound Pith citation observations for arXiv:2508.04195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04195 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:51:32.702267Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:01:52.696984Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:25:02.434510Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7329c7c-d665-42d1-83e9-0297e80d850d · outbound

This paper cites , " * write output.state after.block = add.period write newline.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.419290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.419290Z digest=sha256:6a8fcd89b03c34191a675f8f2615bad82f58b62024a1b8f1f02531e0a60b795b

Observation 72272a94-9074-45d3-9444-56577ac6960d · outbound

This paper cites write newline.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.470752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.470752Z digest=sha256:bd631e038308bca97578be150dd262968a84334636f4eea45dcf212467038248

Observation 001b74c3-9d4b-4ac2-b2fd-de1a1b7f24b0 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.524524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.524524Z digest=sha256:1ed61a782bb944f4bf11a04bf1482f7783867f22c886f408186981b148296da3

Observation bc2d1ba8-8745-4564-bf8f-7f57b24c0aa9 · outbound

This paper cites Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:51:33.246231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:29.579149Z digest=sha256:f8c89b2752afc471987a2eedee357013badd0a10df3ed8e8b19c834a5d3c2307

Observation f7251430-0160-4158-b836-a7496021ba59 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.678062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.678062Z digest=sha256:b374ae91d624a1246d5430ecc82833a7b2b79dbbf3af896c8359acf4083cb535

Observation c2d1e50e-b985-4320-bbcb-8f6a59ffe200 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.737750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.737750Z digest=sha256:e1b95cc923c8e71a48f0aab8d5be1bcfd5fb6c463c4d5ee79d1a920312d36b1e

Observation 9ca561e2-bdbc-49fa-8171-9c771a950e8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.827602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.827602Z digest=sha256:ad985867f9a3cf079a3dbc2487db58fc626ff4410fcd0649451bd1ce8dc54edf

Observation c3fbd9e7-19a0-47af-93b3-4c819a8ee3e5 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.556797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:29.891439Z digest=sha256:2d01931b544262490b5841a5f423ba11801797a86f704ea4b6a1ceb2e9d0144e

Observation ea8f7aa4-6cb6-4a43-a588-7b5dfc2e2fbf · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.957770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.957770Z digest=sha256:b2ea0ceac81b5458c004a81ff432ad53ae7655313a2d767bcdbf6db678a310b4

Observation b679bf31-2d50-4682-94bf-78ffa0bfacc2 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.009070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.009070Z digest=sha256:b12a04d646d777374c39035a52618fd0d9d4c04b186a61b08d707d34389253f6

Observation 15c7e2f3-41c4-4e96-ba2e-5ba58fe90610 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.106597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.106597Z digest=sha256:4e03e4e653f96be0aa334beeae9a15b8f24b1a1b4becc6b280cf050b72271252

Observation 89df5f50-df60-457b-bc39-d774e3dc31b9 · outbound

This paper cites F.; Ellis, D.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations F.; Ellis, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:36.375230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.179764Z digest=sha256:b1b830e825dd8d94be18ffe8d57811a37bc0f067d82dac4a41e9ea050edcf512

Observation 331361af-2510-447b-bd02-e88df15243df · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.216252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.195673Z digest=sha256:8287b388cd6308fc28c71ce3a4488b3fe61b855ca72675064ae402cddc7ae42b

Observation 0deb0d0e-f8ba-4bf9-bad0-278bc699f7b0 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.042699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.296507Z digest=sha256:242ef46c4e477dd6e118f59502f6ac446fd64f28baab715ca928e7fe560756a5

Observation c671d657-7cf7-46fc-a3ff-b7c02cdf8e9e · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.398007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.398007Z digest=sha256:ae4246bafcf5c746b921a720eae6b69a58d53f8a862d1567d29bee856138990f

Observation 4b1c2fe0-0dfc-4965-84d1-e33b4bc65832 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.882856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.527531Z digest=sha256:81aeeeae941fa42e9d5f6cd224800bf2a96b190f349b6b4c2f0f24d09e8e4400

Observation edc42f4b-f738-489b-bf02-57c465a9e932 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.658204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.658204Z digest=sha256:7c73459e9afed8f0b4617fdc937f176751ab366869accfa058b8d90042ea2ddc

Observation ee47a100-26eb-4ef5-ba87-838361fe724a · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.769766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.769766Z digest=sha256:f9799b13fa4e4c4dfeeaf6105d87d681ec652c041d300722794556751bca59de

Observation 3b946c59-8ff3-4d57-96c4-14438a6cac74 · outbound

This paper cites S.; and Aslin, R.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations S.; and Aslin, R

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.727651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.924211Z digest=sha256:0c25b388ded340e6d4b23eab946dbdf2feea63dbd99ab16ae8fae804e63b8fcc

Observation 6f58bd19-c46a-484a-ba33-d99fb23b5c44 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.582626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.118101Z digest=sha256:376a93b5543082ee792d448b1388e955025c872a48efcb1a08af706058bd86d8

Observation c48991fa-e5de-4190-a353-fa8ff95e2e32 · outbound

This paper cites M.; and Nass, C.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations M.; and Nass, C

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.440905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.272438Z digest=sha256:5a8ebb9154cdbefb85b86942978e62115d91129c2f6ac889e5ad018cce25909c

Observation 7d22d4cb-c8e1-4550-a9c7-783d4027daf7 · outbound

This paper cites E.; Rohde, H.; and Corley, M.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations E.; Rohde, H.; and Corley, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.276295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.335779Z digest=sha256:5ac043ce57162062bc8d7512e0f381aae75c1ffd49b55cec14145becd923cf23

Observation 4402f487-40f7-461f-be2d-02e40eaa173d · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.393989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.393989Z digest=sha256:f5bf970dc8302c1fb1d7b1951d3f04a4a3f6c5afb08bef1e869f8db78e00bab5

Observation 90a046bd-2328-42d7-be3b-53404aae77a4 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.131919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.467268Z digest=sha256:156509ec9facb763d53229b73414aee9918d608f00e690bf629cf71e2bfbe38e

Observation 4d568a58-48f6-438a-bb20-7a83a32d2c25 · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.540103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.540103Z digest=sha256:035ac9e7de453af457580928bc89007841815d7d589cbbdb03c5ba947cbd991b

Observation e61ee274-39f3-4365-b648-63172564e471 · outbound

This paper cites M.; Li, G.; and Du, C.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations M.; Li, G.; and Du, C

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:34.968392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.638027Z digest=sha256:2297ac1e5c4bb63d31bbaecb0f49387655c3d032a0e53bcb514209bafd258c02

Observation f1c1672f-6204-4d98-aea6-49563bb65895 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.794508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.749455Z digest=sha256:3909c2fcdce3c32371563b19e04aa0b046b7dd336cd1f610601981d5e0ed9e83

Observation 2ac5f453-ff97-419e-b111-48e6799535e2 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.866776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.866776Z digest=sha256:8baa8c37887daf15ec59fb474551adbfe24ceead13c3398435d8d7cf921e91f3

Observation 350cd4e5-6116-4810-85fa-692f7834aa6d · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.623200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.939646Z digest=sha256:d524a74bf8f8ba7f14dcd88345bcbc049b033112cbf1b27218102e50efc2dc6d

Observation 85795783-f8c7-4e9e-9ce3-5e4595ae444a · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.396959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.025999Z digest=sha256:60e164cabd1a6bb81bd20fb079b22b92ef87e96bb0e370236c960415046075d6

Observation c679a6ea-d81e-483d-a059-f530be6fd47f · outbound

This paper cites A Survey on Neural Speech Synthesis.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations A Survey on Neural Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:32.131382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:32.131382Z digest=sha256:205f4a1e54006190da2bc5c707c19b3c07dec27b322fc224c4316cb28ccbf29f

Observation c2447670-cbd6-4f8c-9fac-3720882b3d2a · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.094456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.233349Z digest=sha256:4d606295643d3e2dc4a116d15b6557ee285e2e2e0e77f6545969331a6c1313be

Observation 24c5b1c2-1b12-4404-9cba-5c97d514214d · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.881745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.333796Z digest=sha256:71ced2a47549051e3f511b1cce2d69d9cb584dc792cd503d5366ac84ffdea079

Observation 701de64b-c30b-4506-bb16-5bd8000d593d · outbound

This paper cites DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:51:32.888098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.388095Z digest=sha256:8b0251c4280fda9cc7b6edcb3f682698823d489c742607236fdc0905757bdd84

Observation fe9c6c71-7323-4996-bf15-f3cd272cfa16 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.665539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.480445Z digest=sha256:39b3abe8aafee380262f5951f25fa7f80bdf12802dca1ae156215c4a05d1621f

Observation e1eb9f8d-3177-42f9-a402-81b86bc9775d · outbound

This paper cites E.; Thakker, M.; Tompkins, D.; Tsai, C.-H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations E.; Thakker, M.; Tompkins, D.; Tsai, C.-H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:33.505371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.557490Z digest=sha256:e3cbedaa9c59d9dfa75d6f6332077f35cbf8b5781175d59d171b5e87d02be051

Observation 352d1629-f1fb-4857-a0c5-0c74c2bd5a9d · outbound

This paper cites Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:32.620801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:32.620801Z digest=sha256:1a816bf90ab5df244dd76743a795cc3e1fcc51624c203eedfeba30c214601c3b

Observation b3f16aa5-609e-407f-8418-baa6bc143a82 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.368415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.702267Z digest=sha256:35100707dc1c15bf7ba8bda05e134bd48c73f6ea1c2de14a0d218bc5d313f19c

Pith citing papers

Observation dc465a96-f3af-4c05-bdc3-c67a12c398db · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.826968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:51e32feb9e473514ed42853af0493acbfb4f9d09d8833a14aa8b6f05c5bc3ab0

Observation 610ee6e7-2250-4543-8895-9e5295d4fd57 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:13.005685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:47d248a4b001d8f88657f8741373f675c3462d8e26cb157a87aeccab82dd828e

Observation ecc84b56-edce-420c-b44d-25ffcba4899c · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.976948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:ed2b85096ae11ffc88d3d9f1b2a56b4480e6b74984f3f7b8a6db576394a21e1e

Observation 16435fef-12da-4d26-819a-82ef45fa8ac9 · inbound

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model cites this paper.

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:07:13.507801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T04:03:27.608638Z digest=sha256:4a8e980388c677e6c096e2ca503e31ab742db141d1fa4f22d145680238293905

Observation 8eb7ae9c-ae45-4db1-b2f8-41d6733a6577 · inbound

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR cites this paper.

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:37:29.604600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T00:29:39.365107Z digest=sha256:763c8f5ac4be16be0907265fabacd050b1fe779ee8443e7ac8f5ae0a051cb57c

Observation 780fd3bc-2880-4093-a9f3-601039e19f43 · inbound

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios cites this paper.

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:25:02.436023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T14:20:24.828073Z digest=sha256:f9d2b78868c25d2a10a6cd130ba5083281aeefdb7c60e6820227f5be693b80f1

Observation d6960b66-d470-4ba4-9ba0-c4a9026b2694 · inbound

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing cites this paper.

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:01:52.696984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:01:52.696984Z digest=sha256:8f431e19febf195b7bdb788b8767cb8980aac6a50ed552b6fb0320b58bbc17da