Pith. sign in

Paper Citation Record · LEDGER

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.16972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16972 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:58.732055Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 749b4bef-0a31-4856-a0bc-f8e3aeb8a981 · outbound

This paper cites online" 'onlinestring :=.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.174990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.174990Z digest=sha256:a07d9f214c091642488b8a498cc0ad41c2e186b61e89ae43373bcf20096d19ac

Observation d5a9b5a6-989a-4d16-b554-ddef40519e90 · outbound

This paper cites write newline.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.259170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.259170Z digest=sha256:3e7d7f552aa9f7ad805d2dfa5af6ed56a5f5a07e6fd259ee83f55e831b7bdc2e

Observation b244d73b-3bb0-4fdb-88d5-9139b5b7122f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.915518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.443183Z digest=sha256:36182efcfc1103220763bed3ea21769356e359ad2bad4372617937ecfa4b75e6

Observation d7b3252d-7557-4296-8eef-e011c7e64fa4 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.815728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.583069Z digest=sha256:949d41b7a07616d090d395128d22bc62ecb6743040218a79e5082cdb01254d6d

Observation 51065def-f9b0-4641-bc23-c1fe90918aba · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.700409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.700409Z digest=sha256:cb56803e96b27f7a8ec316ed6aa4672d1ee58de6e95b1cfdc225ad65e26a1c0c

Observation 2c8ad1b7-942e-4d74-90aa-0c1b497eb2c8 · outbound

This paper cites Synthetic Data from Diffusion Models Improves ImageNet Classification.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Synthetic Data from Diffusion Models Improves ImageNet Classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.847905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.847905Z digest=sha256:bae5ded8d25a636fec0fd87ccfa57df9c225d8a4d10b16b38f57cb21b5b2bf40

Observation 07068995-31a3-45ad-8482-05f6f88c067f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.670409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:53.982528Z digest=sha256:9d217dab4a8116d52a8226278d81966abd06a47405a73ef83c3451c832f8c907

Observation a52c2ee2-3e01-4230-88d4-43cd45c4dc80 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.532208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.154721Z digest=sha256:1ec90bb1c6e55d417ca546fa42c42630859caf3859ccb1e0de9bd8b8ca77972e

Observation 463240d4-c012-4afd-8079-e7ebc5705b0d · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.314575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.314575Z digest=sha256:8e885f6b577cd6ecd618c82fdfe0bb7235e8e9245a8af9a3ea9a560cc8a3f45d

Observation 888acafe-876e-438d-99b9-60c6a4d3cabd · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.427250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.427250Z digest=sha256:e2de7d8a4933d607fb63dba1766c58f79f2754c907c98a40f79a38ad5ad08ef3

Observation 15b3e51d-1b6c-483c-9af9-a277b079a49c · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.383130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.498525Z digest=sha256:29fd2cc82d5a92a9cafae058333253b8b7e83aa7841c05a0e88338dc3ea90363

Observation 35e732a5-8035-43a1-8e9c-66ff798c8af9 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.244325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.566641Z digest=sha256:1aff50ba8ca16f7b24e8571ec0b6fedb470d0600b088a398d66572426c6b68f1

Observation fb412036-97d5-4785-9018-d1fe2474c66e · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.665380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.665380Z digest=sha256:59ac5b2ca1175e52e0f6faf840bbad201f37a21e267cacc47c17b5143f16a09c

Observation 9b65f5b2-3f11-49c6-800e-ef63fd2fe423 · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Towards Robust Speech Representation Learning for Thousands of Languages

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.752647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.752647Z digest=sha256:3c64e9974bca83cf67f0c0222372cd59f456810b5b0a4cbbe459415ba4d08157

Observation ae97234f-0160-422c-af4e-9c536861be1f · outbound

This paper cites Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.842142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.842142Z digest=sha256:c51f59854c2605cef56b8a197b2806f0b469e626979764c7310ce6ec4086d728

Observation 851d97c0-f3eb-4d2c-b8c6-bb67bb7072d4 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.912281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.912281Z digest=sha256:048d8e2bd900928a18429e5e2765b371cb60eddfe320340cb8a7c94b4306d270

Observation f2b799f4-8ca8-453e-9fba-cd4eaae2f49d · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:03.105039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:54.987167Z digest=sha256:92b93ec4f4182506e8c9eded97591661a0d91c4fec730332c2b16fd84dfb0927

Observation 3c0638db-9725-4da7-8d96-67e4b858b084 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.065966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.065966Z digest=sha256:9e48af2b173eb7d1e28411aa36d83a2bd8ebff9afdfe06de718da6706a0385d7

Observation bb08feb4-246b-4e5e-b096-79b083b6457d · outbound

This paper cites CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:55:59.809037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.134893Z digest=sha256:7877846b9e8135bc6f47edcae3b8339fade087cda82564ca956c49fcd8d1b53c

Observation 42604363-4e17-44da-9d12-57971fd4abbc · outbound

This paper cites Understanding Back-Translation at Scale.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Understanding Back-Translation at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.225525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.225525Z digest=sha256:491bf2623fb023dc12f52b023af9b5605ddc408ac8cb70ef52cba0d83826fdee

Observation 614ab78c-b34f-489d-b9a8-3d7585b2a781 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.298928Z digest=sha256:86016b18f490a8593c2c53ac5aadf7cb83abbef642739b444e4418207267d14c

Observation d695a88a-0cf4-4a2e-9a08-59f3d2684eed · outbound

This paper cites Hasegawa-Johnson, Shiyu Chang, and Yang Zhang.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Hasegawa-Johnson, Shiyu Chang, and Yang Zhang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:02.806913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.388526Z digest=sha256:2a6e12806c5e9b5dcf678598237ee8ddb59c5a9b191f8a0b19d7ba341f27bf59

Observation d9558d3f-fb26-44cf-af03-08a9bafb6415 · outbound

This paper cites Textbooks Are All You Need.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Textbooks Are All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.478048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.478048Z digest=sha256:d3241a8dbf15d8a346789c3faca6c880283eb410ec1bcd56d6a21ccd09c2ae18

Observation 373baf00-f961-45f3-b761-91202f53915b · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.684344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.542956Z digest=sha256:f42a9126f87e120f0f604888532cc90453122898a315b159c3de7f2a4802fc97

Observation aa6c3185-dcb8-4aa5-93ae-5dd8727a85ab · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.604407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.604407Z digest=sha256:6e01c5c382b12d05d338833bf3c2223abf811dfa4422b864212578a6ea5ae620

Observation 2a34c5b7-44a9-4d22-90ca-e3fe94a7f1f4 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.482808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.661496Z digest=sha256:db94cd55d3c17b74bdb3f2639c263bd76fa55ead93c85e9fa70cab762be32eb1

Observation edee7b2d-1094-48e3-9e5f-16779cfea925 · outbound

This paper cites Speech Translation with Large Language Models: An Industrial Practice.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Speech Translation with Large Language Models: An Industrial Practice

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.748608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.748608Z digest=sha256:fac86c562195a0d5ca4afb0f72c899fdbdb32fb835f65b8ca8b19b36c112ddca

Observation 57c6908c-707b-4e18-a079-a69d3e6608b0 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.381898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:55.855819Z digest=sha256:c4e4731bf17cb7750df173d8f85a6e076a25cead8adcc4213b8ca9e7c05e1a37

Observation 968f3035-3656-43ed-aa5d-cb8f91ee8ba5 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.933535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.933535Z digest=sha256:5f1cf7dae90ecd9a87a1a44f74b5509a8c3444d31d860a86ae6033cd6c100657

Observation b1a35c0c-15e6-4e3b-931c-20a8f00fe4df · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.230660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.043409Z digest=sha256:a78aa91bdcc1c7786d947bb4dd8edfe607e4020853d31776322f940e262d4366

Observation 704475c2-887a-4818-b96d-c4cba4812de1 · outbound

This paper cites Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.193027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.193027Z digest=sha256:ac644f705e7ea8a73dec5733fee8f5af5c1298f8e60a4bbcd3ec5b90d289e2d4

Observation 3d92d0db-9563-4a93-aa41-6c2dace948a2 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:02.087331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.305477Z digest=sha256:3772f66aaf9ab732c8aef2d0148e71e9edd927c870b56f36882194fef6ea7a5e

Observation 0a8853c1-bc77-4cc2-b259-aa41ef251803 · outbound

This paper cites Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.404903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.404903Z digest=sha256:f2c3e9dd869ab0a83e47e1eb6a1c19b2979cf3edfa92bfc8ea77af06d6ff3646

Observation 48978583-1258-41b0-8653-d9ddca2174a2 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.498003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.498003Z digest=sha256:c3279efe4d5fde0cb1f00fed317dc6c6ca2d9ee634a710cfde6dd8251a6f08de

Observation 875e3c76-392e-4d58-b291-005c49d5375b · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.615438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.615438Z digest=sha256:f2e25d4ac6cf8ec9c3adde2c7613bd4e0ab82a2d72fb79312d7d0b03057f52c4

Observation 99b7891b-7077-4c39-b587-7e6c1e7f1550 · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.665186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.665186Z digest=sha256:de8eb3fe5294c98a5d3122c9a546e02547115cbb0776fade06015df023b6a782

Observation 7489012e-c329-4f5a-8be7-8331d9ac307c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.715907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.715907Z digest=sha256:3b155ab6230102e8fb9821bd9c01e330dc81a60c8dfefb5ecc60e6ce86dfd2c9

Observation eb7f802f-3fe0-4d57-babc-bef3dcafee9f · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.957908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.774284Z digest=sha256:f3f135430dc8f012fc4c1dfbeac291d1132cb574dfd3c9769949157c9fbed85e

Observation 259d6162-e9a0-4a07-bb3d-090db4b295ae · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:01.810253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.858264Z digest=sha256:5198f13f645511798752955906bf81a480a19684e82476b911a4f5430af510d5

Observation 31a069dc-96b6-4896-b176-9adfe6b9f8b1 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.596909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:56.967930Z digest=sha256:ad60453f29c7fa65f7e6ddfc3487529d89fda39121b7d98cc4f30c45d9005c2f

Observation 48d01bbb-d7de-4888-b977-9ccb38332a0e · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.413813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:57.112780Z digest=sha256:cee3dea837c6507d4cd6af00d67bad6af341df0b7ac8c1d7582ee9b4e722c536

Observation fa95a800-6595-417b-af9a-f94e5bea44ef · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.198647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.198647Z digest=sha256:2daad31dd607d915f2042a9456141a1e22c62b0e2cde4b7d1712b10b1719ccbc

Observation 3e8cffa8-06ad-475f-8c14-b3629bb6e103 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.264342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:57.312465Z digest=sha256:69c759050e8935571ed0d9fcfcde70476f0e22bff6f9fc035e3b889320640c1b

Observation f7a78e29-80ab-4c3d-b9aa-f77f7ab8bf71 · outbound

This paper cites StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.433402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.433402Z digest=sha256:468b157a7a5cc936c180b540171a3f7817bf6888b4390e3c8572bc5aed876872

Observation 6b8e309e-b3d5-48c3-8f9d-c4a615853a3d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.537044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.537044Z digest=sha256:cdd45657a6d9790d37be087072ea07c97f952eb8177c436ded28037249d31004

Observation e9b37409-a26b-4b71-919c-2247ad31a3d0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.615225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.615225Z digest=sha256:5c8297f7f51fa4b17749b672c37f0d83a73c3dc2c74e27bf41017620ad7698e4

Observation 1b39a7ae-1242-48e5-a95f-f55e7ee62b94 · outbound

This paper cites Effective Data Augmentation With Diffusion Models.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Effective Data Augmentation With Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.729459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.729459Z digest=sha256:b5694f8bd7e2689aafa2f1ebc6afb3ffe27ce8f833853d4eed25b52209b3faa2

Observation b12b32fd-de54-4556-817b-b2c42d11dca1 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.846781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.846781Z digest=sha256:a2a6e740721e8297d19b97620e7866bb93d01f1b2652ec957bb69e246daf271d

Observation 9dd89f1e-ac5d-4aed-99c1-9262f5d8728a · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.998249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.998249Z digest=sha256:539d7fc9541acabb07b25c973246d048320d05bc189e14ed1694b47fa59e8784

Observation 14196f2b-ee01-4167-89f4-46fe289857e1 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:01.063964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.086002Z digest=sha256:499a758ebc5595aad739a159dd5ff71c181fbef13175477dfcebdd999fd1158e

Observation 0b620978-fab8-4bfd-927c-dcf3d9c6e675 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.913779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.222753Z digest=sha256:3e1bbdb3852c2ed986272e0d30770ddbfcfbcfd7ca823ac950d75aed93be802c

Observation cb572e1c-c837-40e5-b3bf-23355b8c2610 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.730417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.293114Z digest=sha256:63adfa482070f90f1a0d5f8f26a982293b63b9fd369849e5a92839950fbf9786

Observation 298d3a94-989d-42fc-ba0b-9a1891f70374 · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Skywork: A More Open Bilingual Foundation Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:58.395604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:58.395604Z digest=sha256:dfd09507ea6bcc6915fcfbd6e1fef3609b99cca746ff715b407dcc8a1e61a8f7

Observation 6b6f0483-c670-46c6-8a84-3915bb604f78 · outbound

This paper cites Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:58.468368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:58.468368Z digest=sha256:986b84753b63b0f25f71a8aafc3f3edc0017218a18a081a5f4af8f1531917cc9

Observation 230d8c32-385a-4d69-bc79-a1f6c7215863 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.581979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.540112Z digest=sha256:0fb97519f9080b3d88aab129dc73b3524d6dbad8bc6e8ecb9e4ff4ecb52f6601

Observation ed7f5382-0d3b-4d65-a29d-a64967f2ff51 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.411285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.638358Z digest=sha256:d66771bda9572db13038e56057314613ea8e16ff127a384fea1c3d832de089e0

Observation 04e01395-1449-4357-81d7-0b687c7b1a02 · outbound

This paper cites an unresolved cited work.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:00.263139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:55:58.732055Z digest=sha256:d7cb5dd699a7947c366c39848d4c0b624d8cd442a204ecea011adc346b9f2205

Pith citing papers

No inbound Pith citation observations are available.