Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 63 inbound Pith citation observations for arXiv:1904.05862.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:02:44.973306Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:30:08.546246Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ce798540-d1b0-4cbd-b4b1-1e8c5ba8d8f6 · inbound
Vision Transformers Need Registers wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1eeae4cf-6622-43e5-90c0-7fcf9e1b1aad · inbound
Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3b75942-03a6-4b9c-bd97-87e0ddbf016f · inbound
WavChat: A Survey of Spoken Dialogue Models wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 185
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a5bae22-7b35-430a-8a8a-6e90caf29058 · inbound
Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7dd5953-4b0f-491f-82d1-bc933439413b · inbound
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3bf2902f-54f8-42b6-861d-8660cf0724d5 · inbound
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae92db97-c9e1-43cd-a818-8e7bd89ca952 · inbound
TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b6dc154-21c1-4a94-a3e5-ac196ba4ef52 · inbound
Real-time One-Step Diffusion-based Expressive Portrait Videos Generation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c221058-7543-4c6a-859a-f85ec72bcbe4 · inbound
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc3f86b-bba6-49b1-bb64-2770c4ced7f2 · inbound
Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5300cfb3-f06b-46d5-a3b3-1c0f4050d361 · inbound
Tessellated Linear Model for Age Prediction from Voice wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2db48f8-ec86-40b7-834a-c5a81a49ed46 · inbound
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e069ab83-3241-4b8a-bd64-766445ddf959 · inbound
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb4deb4a-4ded-4ddd-91a1-20d164e9dcc0 · inbound
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04c8143-d829-4d61-995b-dd344e649c4c · inbound
Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4e6b38-1910-4a57-a1ed-bb0b929aa685 · inbound
Evaluation of Deep Audio Representations for Hearables wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6f2899-cca5-46c4-9a03-c33486018c95 · inbound
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbe7391-59f9-4613-81fb-6e3e9b2623df · inbound
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780d4055-048f-4c61-9f8a-fdb0d592921b · inbound
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec376700-f940-4ad6-bb8d-14b1b91cd02e · inbound
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c47dc4f-ccf0-487f-9318-a04e7d3f7167 · inbound
Model as Loss: A Self-Consistent Training Paradigm wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a5ef45-374b-4484-983e-4c3fa9ffc8f5 · inbound
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6496d82-3799-4a4e-b019-ed3142d7a88e · inbound
Automatic classification of stop realisation with wav2vec2.0 wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f091b9-a0e6-4a24-993f-f5a1d4d7890b · inbound
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02691a0c-79b2-460b-9721-f69bf7e95692 · inbound
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cd0cc7e-cfb5-4cac-8002-73ea19f1c9e0 · inbound
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbba3595-c4d5-4e38-9e09-6c455e5d1628 · inbound
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eded0f9a-9741-4aa3-b944-1b04c6b765ce · inbound
Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf8f96b-d270-49c2-93b7-92681762c5ff · inbound
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2f536e-d363-4745-ab5f-ffb2d3fc48e7 · inbound
Seeing Voices: Generating A-Roll Video from Audio with Mirage wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41b945c-92e7-4dce-9c77-ab82425a6e3f · inbound
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4311960b-b606-41e4-a1f4-1bd4c7ea2419 · inbound
Manipulated Regions Localization For Partially Deepfake Audio: A Survey wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d44c9b3-5ea3-4053-be17-e2763b175e0f · inbound
A Dataset for Automatic Assessment of TTS Quality in Spanish wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568f8819-1c0c-4d3c-8df3-6da741a4fd4e · inbound
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33d1aa8-4a70-49d7-b9a9-a28f852fa65f · inbound
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827e119e-ffb4-4984-a5d8-e2dd6d862668 · inbound
Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df222587-a596-4f16-9829-02dffc8e538c · inbound
Scaling and Distilling Transformer Models for sEMG wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13b9691-63b7-4200-b309-67be32579770 · inbound
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0dfb18-be61-48e4-a07e-b129e4107d22 · inbound
EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a928eae6-f83b-420b-8800-168703fb3db6 · inbound
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a00d04-42ad-4dd7-a6ad-02046ed1ef14 · inbound
Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd4856e-f8aa-4d72-be37-0229cb13e30a · inbound
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df71453-5bde-4389-8e4f-9818f88521cb · inbound
Entropy-based Coarse and Compressed Semantic Speech Representation Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c737ac1c-8927-4983-8369-8b7f16955cdc · inbound
Contextualized Token Discrimination for Speech Search Query Correction wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8696e58f-7177-4cd2-8f24-63c0c5ab6eed · inbound
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ada0bf-9a1b-410e-9062-db6b3e576a50 · inbound
Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation accda2fa-7009-4344-9ada-51a61be75f66 · inbound
A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10f075cd-643d-4c24-960f-f04cbcc4213c · inbound
The Indra Representation Hypothesis for Multimodal Alignment wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 130decd9-62cf-4431-958e-298031d78e5f · inbound
Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f2f67d08-f403-4fa5-a139-9ed8de2bebdd · inbound
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6eb36e49-f82f-416e-a613-f46fef0c9cf7 · inbound
HighSync: High-Quality Lip Synchronization via Latent Diffusion Models wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8c0ac0c5-0803-4aec-a953-4a25a3e1dc4e · inbound
EGI: A Multimodal Emotional AI Framework for Enhancing Scrum Master Real-time Self-Awareness wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ea22beaf-3b1f-4edf-89d9-0d39535ec478 · inbound
Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12eb3e77-30c5-4fac-bb4d-63cbaf818c0f · inbound
Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a866f1c8-24d9-475d-a7a7-6a4db8a68998 · inbound
Extracting Governing Equations from Latent Dynamics via Multi-View Contrastive Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 24e94bc1-9898-4567-940a-96cebad7cf17 · inbound
Fully Differentiable Neural Forced Alignment via Soft Dynamic Programming wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4e0ca598-ac30-4646-9f0a-4b12dd9ea6cb · inbound
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 93435d25-db93-479d-8e4f-1ab7fdc1df1a · inbound
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3bf9f104-4fb6-4c13-928f-d12585603dc3 · inbound
wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2 wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8cf086c1-9165-45f6-8adb-5028be3f62b0 · inbound
Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d26ce9db-c44b-433a-8ffc-a124754bf397 · inbound
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ce08860b-31a8-4bb5-8558-5a1c77aefc8b · inbound
Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis? wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ceb5d73-c69e-4672-9245-0a033c51e0f7 · inbound
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation wav2vec: Unsupervised Pre-training for Speech Recognition
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.