Pith. sign in

Paper Citation Record · LEDGER

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2101.00390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2101.00390 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:52:53.410512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:42.926035Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5f067a1-5f74-48e5-bc36-15b8f732aaea · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.473790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:36f5d47662efa1ebb462f911bb463c741b99356792d4893bfc033ea81611f758

Observation eee5a22d-e625-4e09-90e3-1b350bbe1090 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:58.097144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:58.097144Z digest=sha256:b17df60672817ae33a0af468b5da9871ae6fa0e4bff4f62ab334df8a8905d828

Observation 0416c9cf-68ee-4316-8720-57bb6deee458 · inbound

Scaling Speech-Text Pre-training with Synthetic Interleaved Data cites this paper.

Scaling Speech-Text Pre-training with Synthetic Interleaved Data VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:40.541338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:02:40.541338Z digest=sha256:846e7028fb9e26ea7d02fbe4639b946c2faf86a553aea7625cbb217eff4f6cd8

Observation 3d464ebc-e6c3-4f03-9102-50a4f111c769 · inbound

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing cites this paper.

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:10.982319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:10.982319Z digest=sha256:bc764f3874efee73580409292e58d5dae7e8ef64416c4f7533b4d58572c04179

Observation b3ea748a-b7de-4c56-8c13-5ab7b14e3b4f · inbound

A Survey on Spoken Italian Datasets and Corpora cites this paper.

A Survey on Spoken Italian Datasets and Corpora VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:59:35.364417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:59:35.364417Z digest=sha256:f1dad0bae071302cc1983efd695543bac9d6588a71ae7c891d2860eb3f57f439

Observation 365fdb58-83d0-4756-99da-81ddc14db160 · inbound

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation cites this paper.

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:16:37.888931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:16:37.888931Z digest=sha256:57b2d860d3f57135fba8e1fb57371f9bdbfe1ec7a66d0a5d9b0b38a87ca583a1

Observation df8b7a56-ea67-465b-8f34-2edfe29b78dc · inbound

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention cites this paper.

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:30:31.517132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-25T08:30:15.011210Z digest=sha256:63a49e01ebf892946060581a07ca27fc2252146af51eaf8b90f99c0037c3fc20

Observation 8c466509-5f86-4db4-80f3-df8a3dd8b06d · inbound

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models cites this paper.

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:49:20.110515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:49:20.110515Z digest=sha256:7ef179bdd49b406bc9f6351e0ab32a8bd639dc2b0bb31895010b20003d13cabd

Observation b4ca6cc9-34e4-439c-b240-26c98d0a3493 · inbound

On the use of Performer and Agent Attention for Spoken Language Identification cites this paper.

On the use of Performer and Agent Attention for Spoken Language Identification VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T17:49:21.497522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:49:21.497522Z digest=sha256:86a18b5aac73ce76cc7278b5328f59a3b02a01b3b523ab02c6b532b4a9da5302

Observation 1c453f19-e2a0-4c06-8853-114f6e5ab342 · inbound

Speech to Speech Translation with Translatotron: A State of the Art Review cites this paper.

Speech to Speech Translation with Translatotron: A State of the Art Review VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:12:07.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:12:07.468458Z digest=sha256:720e75dd35ecee6e8bb8c42c74737fdc780493f1dc276947b55085447885814a

Observation b8dc698b-945b-4e0b-bc4d-6653c2e744d6 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:08.009413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:08.009413Z digest=sha256:bad166192d645aa3b12537357d9f94e768d2ae446bca49f8997a3fbd3cb2ec3e

Observation 444d5cec-fedb-4c44-87c4-6ac9730a923b · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.071481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:0b09913acd012a9b7a07bb318428daf6f28f766598b0ba0ba737258b0d302b2c

Observation 4dbce808-0594-4b3e-8ae8-6d608a832352 · inbound

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities cites this paper.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.410512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.410512Z digest=sha256:5eeb7b059a5fc71640296de63e01429c12e30c40e305cd50a529bc8de70decf1

Observation 339a71f8-d533-4f2b-8879-123ffeacb751 · inbound

Inclusivity of AI Speech in Healthcare: A Decade Look Back cites this paper.

Inclusivity of AI Speech in Healthcare: A Decade Look Back VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:20:25.387876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:20:25.387876Z digest=sha256:c20f6f60b4c96b1d232ba5c3961e773265746cff6ff507e0d761982454ccc117

Observation 3cfacb16-af02-45e6-9163-d43e839a6056 · inbound

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification cites this paper.

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:32.780424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:32.780424Z digest=sha256:8a2ae7fc732e4a92ab77fb7457a512ea52256705aa0602cc4026ea7f6a8145fb

Observation ee506292-b81c-45b4-8312-f35493fea5df · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.237399Z digest=sha256:4ee358a965d0d44a9e6f0842126630ad9403253b2d55da96632294cc13c815aa

Observation b12b32fd-de54-4556-817b-b2c42d11dca1 · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.846781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.846781Z digest=sha256:e3e0e49da9a0a35c533179dd83947f6d39cfbce5138bd0166b37b47610fb54fa

Observation fc9a8ab5-343b-4ba1-ae6c-dedcb0b7b62b · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.668479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.668479Z digest=sha256:74132d6dd8d20524771817e9595248f57a44cb071a8b2bf5ed47afc5bbb178f5

Observation 62cf6209-dec3-431b-98bc-12855fdd7f48 · inbound

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition cites this paper.

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:26.212501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:02:26.212501Z digest=sha256:51d957bd4fb6b0e7278d1f9c57cc19730974619caf14c2e6baff3bbacf0b0be5

Observation 51c5bdbc-2ebc-4811-8a2f-a6b72e0b28e1 · inbound

Unified Semi-Supervised Pipeline for Automatic Speech Recognition cites this paper.

Unified Semi-Supervised Pipeline for Automatic Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:13.888662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:13.888662Z digest=sha256:b8c1eea19d578991854fb068274cc16676b428d7313178cff8eb4348c06213c5

Observation 2ae6d9e7-e12c-4cd5-98cb-8ce49c00d530 · inbound

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching cites this paper.

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:52.578363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:52.578363Z digest=sha256:fa8094b556cb18c73c49c9d5f85a447433d4e81789a2d1b926b9712847bd9b3e

Observation 9d90560f-7ae2-47ac-a5cf-b99fa37eaee7 · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.173525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.173525Z digest=sha256:dfbe3a38e80bc5eed24c5e283cf073e6a663d3ccb8539ffa5e59f29846081dc5

Observation 2e9d6ea3-8a1f-4ed8-8865-22e14f8480c2 · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:18.489814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:18.489814Z digest=sha256:54b62db11f717457d0fa479266a27627814d95b4c226acabed70471957b6df5a

Observation 3ece12cd-30a7-43f6-b941-35aafc10cc84 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.248930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:93542d67a7cfe8d849fb68aaa4c54e9afa791cd6dc9fda5cbbd09518e30fb1ee

Observation 6d9537f6-57df-417c-a1d3-c98e3510fb15 · inbound

On Barriers to Archival Audio Processing cites this paper.

On Barriers to Archival Audio Processing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:08.902551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:14:08.902551Z digest=sha256:04a81b4cceab71c1ebcbe509f572bd017c2de70e2b1995ad0fd3487c0cd1ec9b

Observation 6404bd70-32fa-466e-b654-406b91370c41 · inbound

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications cites this paper.

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:13:09.439905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:13:09.439905Z digest=sha256:11816d473d06eb0a479ad5b8eb229451ae78a7efa09b9a89431a863045aa159a

Observation bf6f6e49-c096-41c8-86db-03287d382039 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.412121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.412121Z digest=sha256:a846ac4ef769f290557bb97ceede38fe5c957559b7a1514f587a236eab553d9c

Observation 1d292078-0891-4d23-92c2-4d347df25a3b · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:25.480535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:25.480535Z digest=sha256:7b2dbd91440b56a1b9f60625082c639fea0179e277a952aede3481b7c3fe1f0f

Observation caa0a545-4f09-45b7-a604-0ee6e3a1e41a · inbound

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model cites this paper.

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:16.274239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:02:16.274239Z digest=sha256:2aa4b99cf08a559d4b8110e301cf0b17c6577a3b59d599a43c2575c29e5387cf

Observation cda344ad-35fd-40b0-8bde-87d3c6790e7e · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.405536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:18225297d8303476a0c52ede13372366893c92cab5b2ba5940e80be50bca0ec5

Observation 512e650d-05d0-4566-8218-f216c51451e2 · inbound

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis cites this paper.

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:17:43.934752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:17:43.934752Z digest=sha256:dcb6c052d0ac6a3f77d3ef9cd921b6079495ba22bf36fdbe8a722485d5bd218f

Observation 9f150e90-e3a9-48ec-90b8-7a5747de2ada · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:02.010823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:02.010823Z digest=sha256:3583cf109b6646b6695b12011468ec2404f54aab8bdd2dfe54962103069156dd

Observation 21093ffe-41f4-445d-8d76-c11f59339e69 · inbound

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition cites this paper.

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-15T00:03:31.986628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:03:31.986628Z digest=sha256:8651c1099421806d6c53e6337deb887f1d449ff886d6b44e18a2201c42bf29ef

Observation fa3ec273-1c9f-4c28-a4d4-a5f4b4cada9b · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.707984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:7edef24b1f7a5319014357ef6a7bfe0da96b8a341ad26d05f02dbaddd3cb4251

Observation c1ff7800-3715-43f2-8791-515dfcb4e63f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.995475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:5695597db7bf3674fcf818137ae51c1babd483a6000a047172eeb2409bf2f5dd

Observation d8f15238-338e-4b4a-bcd5-70bc8f4a5c2d · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.398729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:5313108732e57e65233ee536fd2d6168cfd9d96cc7af42582d6b558136feb4cc

Observation afe61cb0-722c-4aac-b777-01fe72d86fb1 · inbound

A Unified and Reproducible Experimentation Framework for Speech Understanding cites this paper.

A Unified and Reproducible Experimentation Framework for Speech Understanding VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.228406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T21:20:16.428207Z digest=sha256:17d529286202c9ab203e0b628c9599eda3021e1cba2abe86977077afa4fae1ed

Observation 50d2ac01-4d92-4ee5-8c71-936d4a543597 · inbound

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation cites this paper.

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.117785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T06:59:29.020530Z digest=sha256:b2485ad74a581b965369a34cc938beec7d7cd75d227d2d5d99c019ccab6bc54c

Observation 796dd2b9-b815-40d1-9b73-63d8316d7d75 · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:42.927597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:43d42f9b7b28377a72ef0d448260bcfdd1eafb1a6282cf4b9d1e2d46b49cb31c

Observation 2611f189-4158-47dd-b79a-961d6fd9cc4a · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.517800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.517800Z digest=sha256:bbad1a3d91300f3f590d15a163a60ad6b2af1c8d6fe5fbea7f9ab86dc8e74bf1