Pith. sign in

Paper Citation Record · LEDGER

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2505.19931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19931 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:21.978557Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:17.790140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:09:22.202944Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 833af4bd-6ec8-4fc9-9f43-057dfcafb1ab · outbound

This paper cites an unresolved cited work.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:27.136051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:17.710909Z digest=sha256:bff21581337d0c94c83dc6406a17cd786e2793fc7e3402c858a607e949c7aa6f

Observation b2d45a79-0e24-4778-83be-8dfdd4d75c84 · outbound

This paper cites Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:09:22.304038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:17.790140Z digest=sha256:6434fc269fa0c98fe2009aab44605d15822e9d3c5f34ce734b5552f107921fae

Observation d908993a-a4d3-4a7b-aa87-c7b505f39f2d · outbound

This paper cites an unresolved cited work.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:26.894993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:17.913467Z digest=sha256:9dee3530c93d46d44ca9b5c9eda17f4848e3e2841e5a72731f610c5d4a47aff9

Observation b29dd7fc-698b-458b-a56e-ed63be51e2c8 · outbound

This paper cites Setup Baselines.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Setup Baselines

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:26.630572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:18.032752Z digest=sha256:91f38eb409a8f88e98320974817fec76a06a055b6967c06970e9dc1f16bf97c2

Observation 74270a73-3427-4686-ba11-8dfc8b9aa3e5 · outbound

This paper cites an unresolved cited work.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:25.994265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:18.329225Z digest=sha256:d0a1b6fd23a60f524bcb99dc268b028283a29f2b949543fb01cf885cf0918a35

Observation bdc7e683-b8b1-422e-bfda-b94050802db0 · outbound

This paper cites U23B2018 and No.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling U23B2018 and No

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:25.752212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:18.466105Z digest=sha256:66a69c4c7e04573e059cc7ce5769ae7a8842ce666ec02d0fa5abd5a543cfbf3a

Observation 251ef0cc-0a20-44a7-bc65-7bb6071fe80c · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.240224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.240224Z digest=sha256:3301f810fc6ef1793cf501b52bc326b3b9a334bf9f483d5a2f705848dda38d04

Observation 41a2e567-a3f5-478a-b645-78cf8f326abf · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Neural codec language models are zero-shot text to speech synthesizers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:25.505324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:18.543512Z digest=sha256:ed38566bc65d3eeb46367e6153baef3483355fbdb4edd97a8ec0f179947f6a8f

Observation d5fbba8c-c4de-4e24-80ce-2f7dffc1e352 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.662267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.662267Z digest=sha256:b35303e22fb5a5d59ce7c672d1906d6f98e4197562f3fe62224acf90319e1405

Observation 2b4b6fda-367a-4857-a8dd-6ea3b5814775 · outbound

This paper cites V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:25.198101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:18.760763Z digest=sha256:5bb14cb34c15dbeabbcf88b66b44a80818cc2ccf7d34b67af95cb625cba747f5

Observation 8d788c02-3a57-4d3a-9989-b5a4490215d9 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.881600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.881600Z digest=sha256:e50bb1a69188f9cdc9c4c387619fedbd5595eb5bb8c6b2dccd9e051169a1eb96

Observation 684f42d8-25fd-44d4-bc51-97933e6dced9 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Autoregressive Speech Synthesis without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.989987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.989987Z digest=sha256:c83cd191d275ccf111a8bb0c6e15dc31fb5c41fa8acc34657bf1bf4435eeee17

Observation 6984a866-e405-425e-a30e-c7a29aa8b3a2 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.122272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.122272Z digest=sha256:209bb10564d230d93c3452d58c6298131a906a51cec40170a53e8da310d85b8c

Observation f88baa60-07be-4a96-93d4-d4fd0ccd9cdc · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Flashspeech: Efficient zero-shot speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.992868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:20.115305Z digest=sha256:ba74216afc54644f4007708280c3cdd9854c9ee9156bb99c00d54a98cbc80533

Observation 371d11e3-772b-43c2-90e5-66f7d3129b07 · outbound

This paper cites NaturalSpeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling NaturalSpeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.965302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:19.356258Z digest=sha256:31e42a652f090f6446e650d75cd769c40c21baf4f0f3a55633835d98562fb77f

Observation 9305a402-9ae2-4fbf-9798-81d8a363a45f · outbound

This paper cites E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.727077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:19.465221Z digest=sha256:665ddfcf36abd32855a9379841a6600aa3f0c773cdbb1def36024eecc4d86433

Observation 9ba44ab7-e2ce-4b97-b982-f296b1bfdde9 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.616970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.616970Z digest=sha256:3d03131b3c09656c0bbee18e21f610763df1220f7a70f21104145ac39db98e5a

Observation 41799ee3-71cf-4f7b-946d-f5e93b852970 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.712665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.712665Z digest=sha256:c66febfd4ea6025fdeb38f9ebbb7a4427662332ddd7f057acf1eef7d3f7c76f5

Observation 08b4c5b3-1170-45c1-a765-7c3edd8bbc23 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natu- ral and zero-shot speech and singing synthesizers,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Naturalspeech 2: Latent diffusion models are natu- ral and zero-shot speech and singing synthesizers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.491965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:19.865268Z digest=sha256:515265a45f8af62255325e64323e6f2fe6bfe5a0d7c3e14d05c3de78590a218d

Observation 17c5c258-3e93-4b08-94f2-f983a1c4fd47 · outbound

This paper cites Flow matching for generative modeling,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Flow matching for generative modeling,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.256787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:20.019809Z digest=sha256:a95f7953de1840d3c9e92602c2f2c7e980f9546e0e4fcc783ddce0aaf8f48e3e

Observation a8aae7e9-da11-4c45-a8ee-211bf39df1fe · outbound

This paper cites DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.827232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.827232Z digest=sha256:c7f3f966bd544ec4486f0363873e048a0ddc4e05c559f3ed839cc5c9a16affe5

Observation a5bad944-e113-4a06-8a24-54f311e028d1 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.217151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.217151Z digest=sha256:7e4d00f9f59ca79f964d027f5790468de954dc83cc1445087f8f65a034887157

Observation ab173b98-c8d8-45a0-b397-53ad21a74863 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.784637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:20.300038Z digest=sha256:319854257f9b430c242eaa7e6624dd01af83c9bc854a72fb5904b03ee38111d2

Observation e124aefa-74cf-4d60-82b1-0205d922c0a5 · outbound

This paper cites One-step diffusion with distribution matching distillation,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling One-step diffusion with distribution matching distillation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.532183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:20.370668Z digest=sha256:c5c2817e4a239c6b115ce538e261b24675350a9f344536e81924a00d328251b1

Observation 99ebb059-0973-4a04-ba54-396442ffb6dc · outbound

This paper cites Improved distribution matching distillation for fast image synthesis,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Improved distribution matching distillation for fast image synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.388656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:20.489667Z digest=sha256:da77a10931a0a446f44aca50971f90076b2636bf7c8526b44ee1a771cbac48be

Observation 92312d31-4257-4686-8e20-50ce1b56e034 · outbound

This paper cites V oiceflow: Efficient text-to-speech with rectified flow matching,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V oiceflow: Efficient text-to-speech with rectified flow matching,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.206030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:20.624976Z digest=sha256:7eefe1a80f19070825ba7ccf88d80fdc9fa3e582153eb6a5c8e1ebd84a518fcb

Observation c2d5babc-5448-4bc1-812b-4f790f6bf6e7 · outbound

This paper cites All measurements are performed on single NVIDIA RTX 3090 GPU to ensure consistent hardware conditions.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling All measurements are performed on single NVIDIA RTX 3090 GPU to ensure consistent hardware conditions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:26.291774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:18.171011Z digest=sha256:a8f3258bd86674aa82b73ae67e56e264c66e7c5f8d58d774b3ed0a6a709522aa

Observation 409fb9d5-cb23-48db-a668-1044ea7a94e1 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.716820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.716820Z digest=sha256:36bdb89c1d9462b2147fb58c11016a5e89d6fd4c3a21b32f49465af4186e2768

Observation 1c17b967-96a2-4dab-ae22-72df9037102a · outbound

This paper cites Classifier-Free Diffusion Guidance.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Classifier-Free Diffusion Guidance

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.957067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.957067Z digest=sha256:59f4701d4bc11c3f88ca1e392bcc027d5b37a836774b7b4db5bcaca95bb585ef

Observation b74a2ead-2055-4bb1-b258-19768e64983b · outbound

This paper cites Scalable diffusion models with transform- ers,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Scalable diffusion models with transform- ers,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.096404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.096404Z digest=sha256:3da39d264f1b4fba38910db40f7f88d6eb8363a41c80890940be4c2aada59b89

Observation 22b49046-9fab-4464-93d6-aa4854da86f7 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.953147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:21.256740Z digest=sha256:7d5671119d668e999b9085aec97e7b623f6a1766d3d2bbe5a894e05835984488

Observation 63cf9ddb-25b8-49e3-bf24-99942b41f735 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.343004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.343004Z digest=sha256:b25593e72c0f97ca7028f03d6b02859ff9a538e2490ff3f83f7ee8fbe7ccc028

Observation 88557af0-f388-4c21-bd07-5ba4fa4ec765 · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.738335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:21.420951Z digest=sha256:e91d06a1a08e6e98c5c9a099228cb2a35f4e8f04474ccffed0b3dbd3d3f2c581

Observation b3bb9008-d1c6-4c8b-a2d7-d47575f68110 · outbound

This paper cites LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.525712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:21.513116Z digest=sha256:4dd890846b16b546068791debc56cdd70efafeefaf93bf9156d12cb228cb9a02

Observation 34134a67-4d3d-4f83-8ef2-fd2e7af92e61 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Robust speech recognition via large-scale weak supervision,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.635133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.635133Z digest=sha256:9dc6f8f4a4f3650dfcc58f1cad899ff00e7be6b3f31e7ea6d02817edce478fd4

Observation 90899022-b9f2-4eb0-b263-3c169974ec9f · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.735893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.735893Z digest=sha256:52f67e9323d866e61b7434b9d28f2dfe761027b84f65dd0b096069d1b6e93c19

Observation 0a384c2e-ad52-43f6-a61c-cf2f639d9683 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.844478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.844478Z digest=sha256:4ae2345786cb7b1b6c9c1e01a525ce245f9c6a8c954c2444b6c356bb52d3fd93

Observation 298d4cb5-81ce-4c41-a23d-7a84b1156450 · outbound

This paper cites UTMOS: Utokyo-sarulab system for voicemos challenge 2022,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.978557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.978557Z digest=sha256:fd9e43e5528f15759c86f6c7133b65f3fa4bce77114ea29e7ffc3ea830467dbe

Pith citing papers

Observation b2d45a79-0e24-4778-83be-8dfdd4d75c84 · inbound

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling cites this paper.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:09:22.304038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:09:17.790140Z digest=sha256:6434fc269fa0c98fe2009aab44605d15822e9d3c5f34ce734b5552f107921fae