Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:31.716604Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.06566.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:31.716604Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:28.526024Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T19:04:31.829062Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3c7bbe7-8cd5-4329-848c-3a3f824ebf9c · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Humans use auxiliary information, such as spatial and visual cues as well as speaker familiarity, to selectively attend to auditory stimuli [1]
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f5f5636b-395b-42d6-9ec2-ed1836a61fe9 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b9e1ac3-5e6a-4565-b09a-05a92e79abbb · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The architecture comprises an AudioClueNet module, a VideoClueNet module, an embedding combination module, and an extraction network
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2b0a58df-6125-4f35-b39a-2c240e557696 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction 3 was trained using three differ- ent training strategies to study their effect on the model’s robustness
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8f53ced-b569-420d-8012-97098fbd582c · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Model description The basic building block of MTSE system under test is the dual- path recurrent neural network (DPRNN) proposed in [20]
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 43779bf2-5a02-4247-a2cf-197af482189d · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Our initial experimentation showed the models to be sensitive to the normalization layers used in the DNN architecture
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d4496882-b3c2-4409-b3c2-4ff9b4db4cee · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9a11ddb9-7eac-47d7-89ed-e135ddd24952 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Usev: Universal speaker extraction with visual cue,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb7c28ff-8c55-4f15-85d5-d271ef7cfc85 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The cocktail-party problem revisited: Early processing and selection of multi-talker speech
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69977751-296b-4938-9819-bd2c7630f7d7 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Single channel target speaker extraction and recognition with speaker beam,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9d8668d-4ecf-4fa4-8eae-1bb253f96e93 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Improving speaker discrimination of target speech extraction with time-domain speakerbeam,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 089f216d-62a3-4e31-8c87-6a76468e0dd6 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction X-TaSNet: Robust and accu- rate time-domain speaker extraction network,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 589c60fe-fc3c-4f76-9965-eb81eda70371 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction SpEx: Multi-scale time domain speaker extraction network,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 76335c91-0715-4543-8483-70d8986113a5 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Time domain audio visual speech separation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c140bb3-f781-4e02-b8e1-db34f6dc92da · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Muse: Multi-modal target speaker extraction with visual cues,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e867124a-969b-445a-bd4d-3759b391df81 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction A universally- deployable ASR frontend for joint acoustic echo cancellation, speech enhancement, and voice separation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 37e0dac8-7426-44fa-b2d8-c9ddd98fec6c · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multimodal SpeakerBeam: Single channel tar- get speech extraction with audio-visual speaker clues,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2e663395-b6bf-4442-822a-a60e879dae13 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction My lips are con- cealed: Audio-visual speech enhancement through obstruc- tions,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cb1c1913-d5b7-431d-93fe-2c193154737a · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multimodal attention fusion for target speaker extraction,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5b151bd-736d-4985-b74a-7bb033376187 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction An overview of deep-learning-based audio- visual speech enhancement and separation,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6724d629-e7b5-4f52-b762-c08f87cf91be · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Moddrop: Adaptive multi-modal gesture recognition,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8af16198-4809-4e05-88a0-6afd2e435f73 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Modality dropout for im- proved performance-driven talking faces,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7abef295-032d-4e16-965f-2dc171142a5c · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Learnable irrele- vant modality dropout for multimodal action recognition on modality-specific annotated videos,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a63c721-16c5-478e-a033-38feec9b9beb · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction New Insights on Target Speaker Extraction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ebc68b-ad67-43a7-90c9-503d5657ccb0 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multi- stage speaker extraction with utterance and frame-level refer- ence signals,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation facb3b7c-3bab-4864-87d1-b872e9bf72b3 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction TasNet: Time-domain audio sep- aration network for real-time, single-channel speech separa- tion,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72479576-39ad-46d6-89b4-d5b5cc3a851e · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction SDR– half-baked or well done?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09c07bce-6864-4834-8d42-672dee8ab427 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Dual-path rnn: Effi- cient long sequence modeling for time-domain single-channel speech separation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b30c94ea-a70a-4bd7-b763-553e6fd65efe · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Combining residual net- works with lstms for lipreading,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 343025d0-6147-415e-81b8-2125e2ea9722 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Deep residual learning for image recognition,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9272abf-7e09-4bae-990e-c5f33dc78deb · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction LRS3-TED: a large-scale dataset for visual speech recognition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a47bfce-57e9-4d7a-b72c-c28ddb18c00b · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Adam: A method for stochastic op- timization,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5dada267-3769-4780-852e-2c6a11c20b64 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Wavesplit: End-to-end speech separation by speaker clustering,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 003b5539-8b65-403b-9361-20b4f280dca3 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2afe8af-eead-4367-8380-18eb5a4428ea · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Layer Normalization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91be2b0-5ace-46f4-8264-ffdeaf227828 · outbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The inter- and intra-chunk RNNs are realized in the non-causal configuration using bi-directional long short-term mem- ory (LSTM)
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f5f5636b-395b-42d6-9ec2-ed1836a61fe9 · inbound
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.