Pith. sign in

Paper Citation Record · LEDGER

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models

As of 14 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2309.12802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.12802 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T07:01:04.631071Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact3
  • verified fuzzy12
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 610feace-f92d-4548-9371-d9df8e596095 · outbound

This paper cites Top 11 speech recognition applications in 2022.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Top 11 speech recognition applications in 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.303017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:a364cb2bdb093457b9aa600a5df92d3a6e0aff0eaf5a349d7cdd3e44f9b10579

Observation c703f176-8d9f-438c-b1fd-f5906d419a7a · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:04:03.132657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:71aa8c658920477a5a8b00876eab37fb22a6da92420f8ad83331c8ea0335f75a

Observation 9e4bd63a-b0ed-4836-a6fd-0e673fe45443 · outbound

This paper cites Specswap: A simple data augmentation method for end-to-end speech recognition.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Specswap: A simple data augmentation method for end-to-end speech recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.299270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:f46041af23911d3fbd1e917923e6a6ee07a72c10534a28eec33c03236cd04e9d

Observation bbd3b0c7-f9fc-44ae-9c20-cef2e7a55537 · outbound

This paper cites Audio augmentation for speech recognition.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Audio augmentation for speech recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.329426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:f2c1e0f2cdb91f9422060f3b1c965fe4e1bdd2a0b9537e5394fa8f8a7cd0a493

Observation 543840c4-cc7c-4ae1-a5fc-f988fb6dd499 · outbound

This paper cites Text-To-Speech Data Augmentation for Low Resource Speech Recognition.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:04:03.127247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:0f310b47580885da24883bb6309d357d6dfc3a5fc2e52be2449f7f734ea37a6c

Observation 59ff394e-077d-493d-9612-6239e20fb179 · outbound

This paper cites Real-time voice cloning.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Real-time voice cloning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.325933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:1aea2dc5c765c925e0de44ae164abbc633a200a73a54c9b1b7002cef919118ae

Observation 0297bd79-7308-4a16-bcfc-edba285e2bb4 · outbound

This paper cites Transfer learning from speaker verification to multispeaker text-to-speech synthesis.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Transfer learning from speaker verification to multispeaker text-to-speech synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.322547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:b306fd913b43e87380a0cdd48d67f606d07d3dfb2ba4399e88ef764e21f6af56

Observation d246a49d-029a-4992-9dfc-2c206e01ea9f · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Natural tts synthesis by conditioning wavenet on mel spectrogram predictions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.339821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:ff7b754b35a39e4b52da0b421a6934cef1bc3365db1462b0ecc7d15a0a9d9f1a

Observation ae0adcf1-d9e7-4671-b1d1-2d49eabdc2ab · outbound

This paper cites Efficient neural audio synthesis.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Efficient neural audio synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.318782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:ac8a43d447ec2f05b5dffab499d72dcced7645b087313601301c626abf446d0a

Observation 82ec4563-a7b2-428f-8d6f-ab4c8f21d13c · outbound

This paper cites Deep Speech: Scaling up end-to-end speech recognition.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Deep Speech: Scaling up end-to-end speech recognition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:04:03.119378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:7a5cbbfd186ff102795785f4df6bf8dbd0e0eedce69574faac695c385bbe619a

Observation 51b2231d-1ce8-4c35-9f44-00cd3c6a6bd6 · outbound

This paper cites Indian accents- just another version of british english?.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Indian accents- just another version of british english?

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.336376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:b7f285de1834649c484eb1cc46ca292d233408ca8ae7ba67efa5262e9adb52da

Observation 90ac3478-5be7-43bd-a3a9-c6365828dbc7 · outbound

This paper cites Nptel2020 - indian english speech dataset.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Nptel2020 - indian english speech dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.332854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:a3ecf9dd47014529b8c657a755222e47327f839b787949b9efe92c016bb6108b

Observation 13ce744e-dab2-47c1-a8e3-feb069e6927e · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Librispeech: An asr corpus based on public domain audio books

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.314566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:944683d6db674981d3f672d41266b3b0c91e19e4bb99e9364278235bba06b06c

Observation 8b80633c-00ed-4694-bef2-2f23c0bd2f2a · outbound

This paper cites ffmpeg-normalize: Audio normalization for python/ffmpeg.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models ffmpeg-normalize: Audio normalization for python/ffmpeg

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.310492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:53c120883ce57e88fe23e9611fef65405cd00c204d9c904d5a0fef5e1c5ca055

Observation 03d92529-b7a0-4b5e-aac1-91c498b94b68 · outbound

This paper cites Word error rate.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Word error rate

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:04:04.306448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:928b9364e5c98daf218d19e3716c27e3c81ea504602b76cd8d4f1a6abb7094bc

Pith citing papers

No inbound Pith citation observations are available.