Pith. sign in

Paper Citation Record · LEDGER

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2505.19931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19931 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:21.978557Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:17.790140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:09:22.202944Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 833af4bd-6ec8-4fc9-9f43-057dfcafb1ab · outbound

This paper cites an unresolved cited work.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:27.136051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:17.710909Z digest=sha256:757bf70c1c6ebd7916af807842d1ccab84747009653b6bef301985b896847f6a

Observation b2d45a79-0e24-4778-83be-8dfdd4d75c84 · outbound

This paper cites Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:09:22.304038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:17.790140Z digest=sha256:a54697cea78ecee29645e624fd21b249034c28d6a9a8f2778b37cabe677a619a

Observation d908993a-a4d3-4a7b-aa87-c7b505f39f2d · outbound

This paper cites an unresolved cited work.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:26.894993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:17.913467Z digest=sha256:f7c913e453bd50234b73f991dc681a529f19f67575fbce548ca4494b22f58092

Observation b29dd7fc-698b-458b-a56e-ed63be51e2c8 · outbound

This paper cites Setup Baselines.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Setup Baselines

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:26.630572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:18.032752Z digest=sha256:cf38263d51a04c2ec6c71090a8245409c0c3b7d4c17cf6dfed6c9a09711a2766

Observation 74270a73-3427-4686-ba11-8dfc8b9aa3e5 · outbound

This paper cites an unresolved cited work.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:25.994265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:18.329225Z digest=sha256:7954ff9dfdd7a7d9baa0fc58ebd831a39b8c404d3ba3b4b21b3e5c77128078c9

Observation bdc7e683-b8b1-422e-bfda-b94050802db0 · outbound

This paper cites U23B2018 and No.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling U23B2018 and No

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:25.752212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:18.466105Z digest=sha256:718102d9c4f58d4c6daf6937cd95a87af71d8353bec9f3ebc5f79201cf09082e

Observation 251ef0cc-0a20-44a7-bc65-7bb6071fe80c · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.240224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.240224Z digest=sha256:21c052ca862a3eeec9ced1dac199625c8e067f5a717f7ee2b2b1207a91753f52

Observation 41a2e567-a3f5-478a-b645-78cf8f326abf · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Neural codec language models are zero-shot text to speech synthesizers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:25.505324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:18.543512Z digest=sha256:0c71adbb43994ca1924b49069f1deccf74b6b790c9d483276fcd563f7db5c4b7

Observation d5fbba8c-c4de-4e24-80ce-2f7dffc1e352 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.662267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.662267Z digest=sha256:c9fd3debe6258b24c5e9cd11d0e27dd1c8dad0ba2a25d540f7900ac3992360f1

Observation 2b4b6fda-367a-4857-a8dd-6ea3b5814775 · outbound

This paper cites V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:25.198101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:18.760763Z digest=sha256:4a94835f3c1fd0b622c086cf1d9049a67d8ebe891582b93c7a4a8ed9d9614b24

Observation 8d788c02-3a57-4d3a-9989-b5a4490215d9 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.881600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.881600Z digest=sha256:07eaa48ca05f7a89e05514e9f45302f7bf3fe619c12dff846775072b15179802

Observation 684f42d8-25fd-44d4-bc51-97933e6dced9 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Autoregressive Speech Synthesis without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.989987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.989987Z digest=sha256:59aca9123f5fd957bd507b6c53a5813248ac50e13f9acd20ea9000372fe4f7db

Observation 6984a866-e405-425e-a30e-c7a29aa8b3a2 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.122272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.122272Z digest=sha256:fce7706e2cbe9b065e9dc4192bb7ce8f32301a05fd90639687c0b7dbca633877

Observation f88baa60-07be-4a96-93d4-d4fd0ccd9cdc · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Flashspeech: Efficient zero-shot speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.992868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:20.115305Z digest=sha256:d297580eadf86182427d5ae87ee6ee78a230275042a416ff2260a20389ec1a8a

Observation 371d11e3-772b-43c2-90e5-66f7d3129b07 · outbound

This paper cites NaturalSpeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling NaturalSpeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.965302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:19.356258Z digest=sha256:f7773be0b1b2f5855d1c6f696579ebcc18dcd2e02dcd4293f01c661bbbf65144

Observation 9305a402-9ae2-4fbf-9798-81d8a363a45f · outbound

This paper cites E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.727077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:19.465221Z digest=sha256:b70b8ea420da847116e87eed5d1fcdcc4bde8c4de118db0e17abfe07cd8848dc

Observation 9ba44ab7-e2ce-4b97-b982-f296b1bfdde9 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.616970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.616970Z digest=sha256:ea5649df7f1c4b648d03883dbea8a3a5104bbec5d41fca1ab0eb01fb70d3d730

Observation 41799ee3-71cf-4f7b-946d-f5e93b852970 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.712665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.712665Z digest=sha256:0eafa36329bef633f6eaa6c782c21266fa25e7ade7baeec4e1618e133569f244

Observation 08b4c5b3-1170-45c1-a765-7c3edd8bbc23 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natu- ral and zero-shot speech and singing synthesizers,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Naturalspeech 2: Latent diffusion models are natu- ral and zero-shot speech and singing synthesizers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.491965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:19.865268Z digest=sha256:d6b04d4ff757752cade673c9091325d6baa9066c31423f5661955a76c8bcb221

Observation 17c5c258-3e93-4b08-94f2-f983a1c4fd47 · outbound

This paper cites Flow matching for generative modeling,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Flow matching for generative modeling,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:24.256787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:20.019809Z digest=sha256:5c718aa13dd7bd5c4f64694e84e3ecf1c994673c84bdd0f5d2ef5acdffd662b3

Observation a8aae7e9-da11-4c45-a8ee-211bf39df1fe · outbound

This paper cites DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.827232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.827232Z digest=sha256:7c08194658341460365007a31b281a52c3c3d33ac429d49e2f59621628700cb1

Observation a5bad944-e113-4a06-8a24-54f311e028d1 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.217151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.217151Z digest=sha256:1670f3abd7b6710ede16e2fed31e12fa40ca1dfd4be438d4c41774f4c9ca2af5

Observation ab173b98-c8d8-45a0-b397-53ad21a74863 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.784637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:20.300038Z digest=sha256:f35bd0c07963ac2a9d6498274e9b7411c81c00050e18c776f6db910d2a5f9241

Observation e124aefa-74cf-4d60-82b1-0205d922c0a5 · outbound

This paper cites One-step diffusion with distribution matching distillation,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling One-step diffusion with distribution matching distillation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.532183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:20.370668Z digest=sha256:73e654bcb0a17cec9f86dc5a52e7fb4c1aacfa9f2f5a787b8d5e3f57a69d458b

Observation 99ebb059-0973-4a04-ba54-396442ffb6dc · outbound

This paper cites Improved distribution matching distillation for fast image synthesis,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Improved distribution matching distillation for fast image synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.388656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:20.489667Z digest=sha256:bfe13a00be0c314b3781e1106c5c9a1e3b3759b3ef2e8ec9d0c01249d9945f4e

Observation 92312d31-4257-4686-8e20-50ce1b56e034 · outbound

This paper cites V oiceflow: Efficient text-to-speech with rectified flow matching,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V oiceflow: Efficient text-to-speech with rectified flow matching,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:23.206030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:20.624976Z digest=sha256:2862eae602400f90098241ed9ef4d325f85cbccda3a114e8ab93129a1974bbae

Observation c2d5babc-5448-4bc1-812b-4f790f6bf6e7 · outbound

This paper cites All measurements are performed on single NVIDIA RTX 3090 GPU to ensure consistent hardware conditions.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling All measurements are performed on single NVIDIA RTX 3090 GPU to ensure consistent hardware conditions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:26.291774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:18.171011Z digest=sha256:9fc0ba8d9e161d77cb68933ae788ec9ee5f6b5c8368f82437f3a02b00ac795a1

Observation 409fb9d5-cb23-48db-a668-1044ea7a94e1 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.716820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.716820Z digest=sha256:630a98773f2c5735106f5df93e5faa66d548d87f7755486350ea3073b2b8af63

Observation 1c17b967-96a2-4dab-ae22-72df9037102a · outbound

This paper cites Classifier-Free Diffusion Guidance.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Classifier-Free Diffusion Guidance

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.957067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.957067Z digest=sha256:d7a4a5e8febe21dbba2b5afa5c2890b85348aca238d99e960d67d94e4463e54c

Observation b74a2ead-2055-4bb1-b258-19768e64983b · outbound

This paper cites Scalable diffusion models with transform- ers,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Scalable diffusion models with transform- ers,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.096404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.096404Z digest=sha256:97f010a0974e2650d02236fa0fed7c09b8845610c892952f9995f4ad49f459ad

Observation 22b49046-9fab-4464-93d6-aa4854da86f7 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.953147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:21.256740Z digest=sha256:defa3821e33acb51c445873f22ae1bc70da77918501c8a0ffda158a263535adf

Observation 63cf9ddb-25b8-49e3-bf24-99942b41f735 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.343004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.343004Z digest=sha256:68e8333f15a8363859a18cc2f788939030148d8814ad04139ebff4a74208b48d

Observation 88557af0-f388-4c21-bd07-5ba4fa4ec765 · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.738335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:21.420951Z digest=sha256:b28b78b1387d3f9d228825233ac150a1f66d023d5d721d99a8fbace82e75a85e

Observation b3bb9008-d1c6-4c8b-a2d7-d47575f68110 · outbound

This paper cites LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:22.525712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:21.513116Z digest=sha256:4dab768f0d7f53115fdc5040bff90ae023e299d9c24325aa5f2a283cfdac7637

Observation 34134a67-4d3d-4f83-8ef2-fd2e7af92e61 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Robust speech recognition via large-scale weak supervision,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.635133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.635133Z digest=sha256:abbb398832c3bc9ea11303a28f0f09d3e05eebdcda6ea6e236b447987ebf793a

Observation 90899022-b9f2-4eb0-b263-3c169974ec9f · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.735893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.735893Z digest=sha256:e296f5514addaa70f1bccea4da474db46d217b72c1332da3ec61fc645c866e8f

Observation 0a384c2e-ad52-43f6-a61c-cf2f639d9683 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.844478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.844478Z digest=sha256:0e657d0374e0282967f0311d32b7e3f9aa97b9550674b5a69467b2cf1c4f1c06

Observation 298d4cb5-81ce-4c41-a23d-7a84b1156450 · outbound

This paper cites UTMOS: Utokyo-sarulab system for voicemos challenge 2022,.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:21.978557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:21.978557Z digest=sha256:c2749c6a4ab3b581e6f4564dda012be2e3b3ff3a7351d256abff9d715ba7e27d

Pith citing papers

Observation b2d45a79-0e24-4778-83be-8dfdd4d75c84 · inbound

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling cites this paper.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:09:22.304038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:09:17.790140Z digest=sha256:a54697cea78ecee29645e624fd21b249034c28d6a9a8f2778b37cabe677a619a