Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.951657Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2507.21448.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.951657Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.733134Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T14:25:47.176963Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d106859a-c6e3-4b29-85e9-e2332ce048a8 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 28cb8503-5e9d-4c82-9740-8c72d3551bec · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ee473f-0af7-45fd-a927-7e9eec61fc6a · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Datasets We use V oxCeleb2 as the speech dataset and MUSAN for noise and music
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe688ff9-3a4b-415d-ad9d-7da2dd460b0b · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 043a9f23-38cb-400b-b8aa-b8a7632bcb31 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations We systemat- ically analyze how visual embeddings from audio-visual speech recognition (A VSR) and active speaker detection (ASD) im- pact A VSE performance
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cc567b3-95f6-49d6-affb-f61ab3fbe687 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Funda- mentals, present and future perspectives of speech enhancement,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6603e162-ab9a-48c5-87de-ad0fe3314878 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a2b8f46e-57a4-4bc8-b7c1-2d3e777d5d76 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations FlowA VSE: Effi- cient Audio-Visual Speech Enhancement with Conditional Flow Matching,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f16c5e63-8812-4029-b90c-e8327b680cd1 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Personalized speech enhancement: new models and Comprehensive evaluation,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec41fe69-5a12-4ac1-9fa3-8f61a6135da1 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual End-to-End Speech Enhancement,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93e19642-7f0d-45c4-b5c1-82a5c30dd41d · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 41dcc2f1-e283-4f1b-bfaf-dba2194b384e · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations The Conversation: Deep Audio-Visual Speech Enhancement,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71d3d07a-a00f-4cdc-ad8b-52e049c05e8c · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Improved Lite Audio- Visual Speech Enhancement,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d7ade20-feb2-4ca5-a041-e58855dbae5b · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Lite Audio- Visual Speech Enhancement,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 36ccab49-88be-4571-84f6-0b66294fdd9f · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations End-to-end Audio-visual Speech Recognition with Conformers,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff992ea9-cec4-4d7a-8524-b16de94de606 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58298bd1-13a6-4b4c-8c64-1303f715c0a7 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations A Novel Real-Time, Lightweight Chaotic-Encryption Scheme for Next- Generation Audio-Visual Hearing Aids,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bfb62780-2c0f-496b-9138-d95029e2e822 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Lip- Reading Driven Deep Learning Approach for Speech Enhance- ment,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 456aca05-6cfe-4c8c-b3f7-9887d564587c · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-visual speech enhancement using deep neural networks,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation efe08a43-f7f2-46c9-977a-4fb752add33f · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a7608f3b-cb83-422c-a4a2-ef088fc628f1 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Evaluating Audiovisual Source Separation in the Context of Video Conferencing,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb2528b1-7047-4d06-95f7-4707018a963d · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations CochleaNet: A robust language-independent audio-visual model for real-time speech enhancement,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dbbd2a9a-9b4c-4af3-b48e-b3414632f914 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations RT-LA-V ocE: Real- Time Low-SNR Audio-Visual Speech Enhancement,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f9666af7-49c0-4b77-9739-43f790274900 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Exploring Tradeoffs in Models for Low-Latency Speech Enhancement,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 461f0d2d-4cb5-4788-ab00-c1ca4ef1c00a · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1b86e74-3ec7-4364-90d5-71c309b90f24 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations On The Compensation Between Magnitude and Phase in Speech Separation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35634989-c821-4f5d-9702-041a675255ac · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Learn- ing Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4be04175-9aba-4d25-8bb7-4c1f9d31a392 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Robust Self-Supervised Audio-Visual Speech Recognition,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c83f124f-531d-47be-bf7d-6ed43c7bc8ae · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Visual Speech Recognition for Multiple Languages in the Wild
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4299ee8-9644-4ef2-b76c-783a57c86573 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8efe7a-2f6a-453d-bc3b-1a8a4c1d2a22 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations LoCoNet: Long-Short Context Network for Active Speaker Detection,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 623afb87-27a1-420f-b9e9-92d3d3dd2098 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Deep Contextualized Word Representa- tions,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7ad79254-0c6a-4a76-add1-0af98ccf5185 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Few- shot Image Classification: Just Use a Library of Pre-trained Fea- ture Extractors and a Simple Classifier,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6cdf9559-167c-4aa0-9035-ae8f6ba41f81 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Music auto-tagging in the long tail: A few-shot approach,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7521eefd-a268-4a34-a654-eadff6bd0025 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Contextual String Em- beddings for Sequence Labeling,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 736da3cf-bdca-444d-8d34-8482e8162dea · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations DeepFilterNet: A Low Complexity Speech Enhancement Frame- work for Full-Band Audio based on Deep Filtering,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a77b55c-e7fc-4f21-b061-7bf99b65d384 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Perceptual eval- uation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 791b7492-0965-4f57-8f69-f1fd60cf026d · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations SDR – Half-baked or Well Done?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cec44eee-4d81-4e17-a0fe-17ad826ce9d1 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations An Algorithm for Predicting the In- telligibility of Speech Masked by Modulated Noise Maskers,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587209d4-2447-4d2d-ab72-2a0335df526d · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations ViSpeR: Multilingual Audio-Visual Speech Recognition,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c3a8019-162a-48d6-b28a-553cc93946d2 · outbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 28cb8503-5e9d-4c82-9740-8c72d3551bec · inbound
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7317e6e-dbb0-42c4-8761-97e6f89fb4d3 · inbound
FSD50K-Solo: Automated Curation of Single-Source Sound Events Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.