Pith. sign in

Paper Citation Record · LEDGER

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.16972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16972 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:58.732055Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 749b4bef-0a31-4856-a0bc-f8e3aeb8a981 · outbound

This paper cites online" 'onlinestring :=.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.174990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.174990Z digest=sha256:7f47457579b64bfe6205da5b53e491e407a87b0d72e39555a92ff37184d8534f

Observation d5a9b5a6-989a-4d16-b554-ddef40519e90 · outbound

This paper cites write newline.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.259170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.259170Z digest=sha256:43afa06c50a7aa3bda90ff64c3efb5d30b2c773728dab9a88ca3043cd72f534f

Observation b244d73b-3bb0-4fdb-88d5-9139b5b7122f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.915518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.443183Z digest=sha256:bb01f797b8ccbbbd9c30fe2231556f1044632aa8511eecd404c0e48a57eae737

Observation d7b3252d-7557-4296-8eef-e011c7e64fa4 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.815728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.583069Z digest=sha256:6638dee6ef3cc01d7413644a130c17fc2a580b914b4e13c6ea12777d8489a9fe

Observation 51065def-f9b0-4641-bc23-c1fe90918aba · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.700409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.700409Z digest=sha256:e001f148c3cdfa6f0a705047e3c25a256ed74a305e2c4a76ba781fb57c8745ac

Observation 2c8ad1b7-942e-4d74-90aa-0c1b497eb2c8 · outbound

This paper cites Synthetic Data from Diffusion Models Improves ImageNet Classification.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Synthetic Data from Diffusion Models Improves ImageNet Classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.847905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.847905Z digest=sha256:9f8b777619c26cbef640e6c098f7c1a33246aa9b5376b2e006876fe0bdde699b

Observation 07068995-31a3-45ad-8482-05f6f88c067f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.670409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.982528Z digest=sha256:599ba5b034f2d594d8c04e548209c5799cf9406a6232ba0f575625bd6c27a606

Observation a52c2ee2-3e01-4230-88d4-43cd45c4dc80 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.532208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.154721Z digest=sha256:9a21c650ced843d414d8c4f0a209427a12e1c881a6165f3c6c26165e90770bf5

Observation 463240d4-c012-4afd-8079-e7ebc5705b0d · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.314575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.314575Z digest=sha256:723957e8d6c75eb6ab53893b142e058013d90c626ae82ef69c693687e89e6ec2

Observation 888acafe-876e-438d-99b9-60c6a4d3cabd · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.427250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.427250Z digest=sha256:96ce424e8550c2d53929801a6a0224c7e711604c0c08295309f6c3d973533402

Observation 15b3e51d-1b6c-483c-9af9-a277b079a49c · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.383130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.498525Z digest=sha256:5c15e6df5121a8d9178cbd7fa6e4d992d77ae6e1b0f6269d3e88dc47f920367d

Observation 35e732a5-8035-43a1-8e9c-66ff798c8af9 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.244325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.566641Z digest=sha256:7031c900faf6b583a7acd8ed88a19628dc06bcee55ae37c44063b9186405df06

Observation fb412036-97d5-4785-9018-d1fe2474c66e · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.665380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.665380Z digest=sha256:be7d084138a4403d59f6aff7e388f178a363f8957aa2c4d58ff4607c70c96e27

Observation 9b65f5b2-3f11-49c6-800e-ef63fd2fe423 · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Towards Robust Speech Representation Learning for Thousands of Languages

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.752647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.752647Z digest=sha256:8a518cee50cb19882a8888bea32db9940aa1ff74c2e37e3a8bb7399cd8b950ab

Observation ae97234f-0160-422c-af4e-9c536861be1f · outbound

This paper cites Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.842142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.842142Z digest=sha256:b68e1f02053a1441114083de239cb5f6e8ff4b9a54f4062cfdfc1bef89f718f9

Observation 851d97c0-f3eb-4d2c-b8c6-bb67bb7072d4 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.912281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.912281Z digest=sha256:68331eb202338245e6bebd9df5e6b629fbbfcf0e700ce2ecf07acc40ed93dbd4

Observation f2b799f4-8ca8-453e-9fba-cd4eaae2f49d · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.105039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.987167Z digest=sha256:773e3024aa35466183e1bca1ba317982b380549878d8e7794bac9aab766f04ed

Observation 3c0638db-9725-4da7-8d96-67e4b858b084 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.065966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.065966Z digest=sha256:5d5c0f02d2c870952da03151e74e86675a2a30573cc4ecf40f2637430a8f49f9

Observation bb08feb4-246b-4e5e-b096-79b083b6457d · outbound

This paper cites CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:55:59.809037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.134893Z digest=sha256:8c14788b6de0765692cba957ef92d938729d2b15e650d266defd9a8a7f10d1de

Observation 42604363-4e17-44da-9d12-57971fd4abbc · outbound

This paper cites Understanding Back-Translation at Scale.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Understanding Back-Translation at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.225525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.225525Z digest=sha256:409cd8c94091d5a50112f535a972188908cefce4d176466f1f05706cf2f83a89

Observation 614ab78c-b34f-489d-b9a8-3d7585b2a781 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.298928Z digest=sha256:52fde101eba6d7a3056804d33030a973efb1abbf4d3505140aae1b6adeedbb8a

Observation d695a88a-0cf4-4a2e-9a08-59f3d2684eed · outbound

This paper cites Hasegawa-Johnson, Shiyu Chang, and Yang Zhang.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Hasegawa-Johnson, Shiyu Chang, and Yang Zhang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:02.806913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.388526Z digest=sha256:84b6d000d662a05a2fa06da674a8864aaf06b58c1a421af48795ff00e35d5640

Observation d9558d3f-fb26-44cf-af03-08a9bafb6415 · outbound

This paper cites Textbooks Are All You Need.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Textbooks Are All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.478048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.478048Z digest=sha256:bc4ea28c7d62b5f8bdf0c82fd5c082b1048552a73abc89922d5a48b346648afb

Observation 373baf00-f961-45f3-b761-91202f53915b · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.684344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.542956Z digest=sha256:8753e7ec43b1f54b18eec65aab7fd65a5acffdba5df40a7776b3bedf2581db3f

Observation aa6c3185-dcb8-4aa5-93ae-5dd8727a85ab · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.604407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.604407Z digest=sha256:d9513cfeaf1ca40a892951a5b830e0bfcfd73663320502747e426eb76cb78c54

Observation 2a34c5b7-44a9-4d22-90ca-e3fe94a7f1f4 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.482808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.661496Z digest=sha256:c648390497b69c49fcdb5a9b1ed46864c16181aec41537edb972dba068cc2c6e

Observation edee7b2d-1094-48e3-9e5f-16779cfea925 · outbound

This paper cites Speech Translation with Large Language Models: An Industrial Practice.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Speech Translation with Large Language Models: An Industrial Practice

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.748608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.748608Z digest=sha256:271cf459b3f475cbabebe2d4c71179ff26a2829d326c18394533a03db369bbc3

Observation 57c6908c-707b-4e18-a079-a69d3e6608b0 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.381898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.855819Z digest=sha256:2bac88a6f9c99413807a4c7578b6f652cbe30ef7d362dcaaef3a5344edf2e8dd

Observation 968f3035-3656-43ed-aa5d-cb8f91ee8ba5 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.933535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.933535Z digest=sha256:69080dc2dc761e182891083b35cfa3dd4211c987c792c4af697bbe6f6930acb0

Observation b1a35c0c-15e6-4e3b-931c-20a8f00fe4df · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.230660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.043409Z digest=sha256:d6d440d0a016a855c137f8ec36bafa77c2e0c48e709fc4ec1d792b3b7efd781a

Observation 704475c2-887a-4818-b96d-c4cba4812de1 · outbound

This paper cites Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.193027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.193027Z digest=sha256:a18179d18d96e365305dbc0fc2b048af3a1b8f47c2be66fe0c0e91990563694c

Observation 3d92d0db-9563-4a93-aa41-6c2dace948a2 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.087331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.305477Z digest=sha256:cd2a20741afd2868524598577766f795d48fa05cb0c81470d7210721dd0e7dbf

Observation 0a8853c1-bc77-4cc2-b259-aa41ef251803 · outbound

This paper cites Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.404903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.404903Z digest=sha256:cf008bf22f901806eca11c2f6f6c29a02832d96e3e8a5192fac8e1ce057086f6

Observation 48978583-1258-41b0-8653-d9ddca2174a2 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.498003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.498003Z digest=sha256:9e020b46439f7486ff141bf2b14e3d0230938262c98310f44954b855e0b7b463

Observation 875e3c76-392e-4d58-b291-005c49d5375b · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.615438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.615438Z digest=sha256:d5320b57bd56deb6a56d9b2166bd4450699612ad1a083d5a5d037e6cb9a7ca37

Observation 99b7891b-7077-4c39-b587-7e6c1e7f1550 · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.665186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.665186Z digest=sha256:310742671da71e8abbda286c0ec2ffe71b948fde1e42e26ad094d5489488d61e

Observation 7489012e-c329-4f5a-8be7-8331d9ac307c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.715907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.715907Z digest=sha256:fd06c935ade6d9d1c2e3c854403d9e0ebaa7bd33614e86569752ec02ac8ed0f7

Observation eb7f802f-3fe0-4d57-babc-bef3dcafee9f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.957908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.774284Z digest=sha256:f6dfc016e12c44ec400ad3a8fddb6cc4fafd29918d5eeff7a3a5e97caec1766f

Observation 259d6162-e9a0-4a07-bb3d-090db4b295ae · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:01.810253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.858264Z digest=sha256:5eedc1a78a6da4876a050feec104b9af6cd73017842887a920410dff964c9b84

Observation 31a069dc-96b6-4896-b176-9adfe6b9f8b1 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.596909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.967930Z digest=sha256:c206300cf620a2784345ef4b7e27623bc604637d31daeaf793487d04996b13a0

Observation 48d01bbb-d7de-4888-b977-9ccb38332a0e · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.413813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:57.112780Z digest=sha256:d41986aea23cb775180b71bbde765905841a494a1a8ba757d5fc68083ee8be6f

Observation fa95a800-6595-417b-af9a-f94e5bea44ef · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.198647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.198647Z digest=sha256:8d80ffff0c747bd7607f272ce800a097ce455e9af43411eab96204f146be3f00

Observation 3e8cffa8-06ad-475f-8c14-b3629bb6e103 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.264342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:57.312465Z digest=sha256:23574632a9a2076ad850fcc603bdcda8b58aa74c137cbfa5925970d1ec64a338

Observation f7a78e29-80ab-4c3d-b9aa-f77f7ab8bf71 · outbound

This paper cites StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.433402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.433402Z digest=sha256:0efefc59148fe97571668c80f335fd9f16d45183cfdc28ab6f7fe0f43b9d01fa

Observation 6b8e309e-b3d5-48c3-8f9d-c4a615853a3d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.537044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.537044Z digest=sha256:0e0b4819bcecae6ef92d846b4d2560b7e9479bcb83626cb5ec52e43b984291d0

Observation e9b37409-a26b-4b71-919c-2247ad31a3d0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.615225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.615225Z digest=sha256:42a75019088f0ec16ddb631fc68392e32b566b294f4ac700d4d99f2db9bcc54a

Observation 1b39a7ae-1242-48e5-a95f-f55e7ee62b94 · outbound

This paper cites Effective Data Augmentation With Diffusion Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Effective Data Augmentation With Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.729459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.729459Z digest=sha256:eb23818e747517d98a40f79d156a7d3abab4c8ed2ad4b7681043e5a0d25a6e65

Observation b12b32fd-de54-4556-817b-b2c42d11dca1 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.846781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.846781Z digest=sha256:c27180fcc451e0ba0e62b47e2f912f03f60daf4ee12daba1b57e643b3d32def7

Observation 9dd89f1e-ac5d-4aed-99c1-9262f5d8728a · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.998249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.998249Z digest=sha256:bd6ff78d7b6f5bab9c3ca8f5255d2d95edc6e89d097deefe1ef9a215328465c4

Observation 14196f2b-ee01-4167-89f4-46fe289857e1 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.063964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.086002Z digest=sha256:da1ac64af1483b701ddb9e1e379299d9d6bdb9e2a06f78d158d01a4299009067

Observation 0b620978-fab8-4bfd-927c-dcf3d9c6e675 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.913779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.222753Z digest=sha256:e4009032f3f4e029ed041184fa6dc0a9c7cf6c7877e7a8de9e409f6a71e889a1

Observation cb572e1c-c837-40e5-b3bf-23355b8c2610 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.730417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.293114Z digest=sha256:9c5277e10cdf00612024531b0507bbbd2c1bb7a4d6c64aef02134f29c53504a7

Observation 298d3a94-989d-42fc-ba0b-9a1891f70374 · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Skywork: A More Open Bilingual Foundation Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:58.395604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:58.395604Z digest=sha256:ae3ee2a443cb8588238211839dd1fb024514bf29ad0a1a7678f8eaa02d9a20bf

Observation 6b6f0483-c670-46c6-8a84-3915bb604f78 · outbound

This paper cites Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:58.468368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:58.468368Z digest=sha256:1cd44e99b09b8b1035f0b6ac6f66bf4921c4fe1380cafa453aa841835b0dcdd4

Observation 230d8c32-385a-4d69-bc79-a1f6c7215863 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.581979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.540112Z digest=sha256:cbaa83c6bb2c6a76d622fbe505ccf6c1483fb1e9b3b57c0a4106fa8e4065f615

Observation ed7f5382-0d3b-4d65-a29d-a64967f2ff51 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.411285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.638358Z digest=sha256:271558cb2e80c104f045ebac79971ca2c35e7fbde786382bdafc0f80cf468c19

Observation 04e01395-1449-4357-81d7-0b687c7b1a02 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.263139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.732055Z digest=sha256:bd47e7f9191b0b834c398456797f14bec35c9600ff7104d39a1c8062673f1de2

Pith citing papers

No inbound Pith citation observations are available.