Pith. sign in

Paper Citation Record · LEDGER

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.06566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06566 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:31.716604Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:28.526024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:04:31.829062Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3c7bbe7-8cd5-4329-848c-3a3f824ebf9c · outbound

This paper cites Humans use auxiliary information, such as spatial and visual cues as well as speaker familiarity, to selectively attend to auditory stimuli [1].

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Humans use auxiliary information, such as spatial and visual cues as well as speaker familiarity, to selectively attend to auditory stimuli [1]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:36.275780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.438780Z digest=sha256:36ed8370949cc14b970c3456519632e9c6a638e64885cf36ae532130e662ada6

Observation f5f5636b-395b-42d6-9ec2-ed1836a61fe9 · outbound

This paper cites Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:04:31.889166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.526024Z digest=sha256:16712a8831550e572e190490778108d09e930e6d4d6a17cacf5067975c0ce8b9

Observation 4b9e1ac3-5e6a-4565-b09a-05a92e79abbb · outbound

This paper cites The architecture comprises an AudioClueNet module, a VideoClueNet module, an embedding combination module, and an extraction network.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The architecture comprises an AudioClueNet module, a VideoClueNet module, an embedding combination module, and an extraction network

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:36.013691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.619912Z digest=sha256:16478e059b53b005f85f41a5631d9ba9174e9d75cc753c3dcf21292cfaf363cb

Observation 2b0a58df-6125-4f35-b39a-2c240e557696 · outbound

This paper cites 3 was trained using three differ- ent training strategies to study their effect on the model’s robustness.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction 3 was trained using three differ- ent training strategies to study their effect on the model’s robustness

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:35.798746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.736738Z digest=sha256:43d88bdb5a434b524c861470366923588875247bad7811e3ea6f8bd70347590d

Observation d8f53ced-b569-420d-8012-97098fbd582c · outbound

This paper cites Model description The basic building block of MTSE system under test is the dual- path recurrent neural network (DPRNN) proposed in [20].

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Model description The basic building block of MTSE system under test is the dual- path recurrent neural network (DPRNN) proposed in [20]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:35.558471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.858366Z digest=sha256:3e09af48b50a3a7747c6743dfcddcd54e1bf7a89cfed8871a85729a83099d60e

Observation 43779bf2-5a02-4247-a2cf-197af482189d · outbound

This paper cites Our initial experimentation showed the models to be sensitive to the normalization layers used in the DNN architecture.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Our initial experimentation showed the models to be sensitive to the normalization layers used in the DNN architecture

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T19:04:35.119973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.080203Z digest=sha256:8ea2fce9849dc9398772104f9c960aa966b93db5797dd83ec4215477a863f15f

Observation d4496882-b3c2-4409-b3c2-4ff9b4db4cee · outbound

This paper cites an unresolved cited work.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:04:34.999737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.182452Z digest=sha256:d34e6a835d8a7d4b9dce6d14f3ffd3497279a27feb6766b13904839328ea917f

Observation 9a11ddb9-7eac-47d7-89ed-e135ddd24952 · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Usev: Universal speaker extraction with visual cue,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.204021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.965218Z digest=sha256:a72a52f06d4d7bbde04e3e6351fe7edecbe7fafa9a28f39a6d2488e44ff9a433

Observation eb7c28ff-8c55-4f15-85d5-d271ef7cfc85 · outbound

This paper cites The cocktail-party problem revisited: Early processing and selection of multi-talker speech.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The cocktail-party problem revisited: Early processing and selection of multi-talker speech

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.896402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.282529Z digest=sha256:5f463b5e616f0fc5dd7702928f34f4953289748b0e5d6ac919988005aa51357c

Observation 69977751-296b-4938-9819-bd2c7630f7d7 · outbound

This paper cites Single channel target speaker extraction and recognition with speaker beam,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Single channel target speaker extraction and recognition with speaker beam,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.794652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.364193Z digest=sha256:4415a30d57aa0cea39ae063317b1cae877a8a628d9b1f3838ee81b603cc01300

Observation b9d8668d-4ecf-4fa4-8eae-1bb253f96e93 · outbound

This paper cites Improving speaker discrimination of target speech extraction with time-domain speakerbeam,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Improving speaker discrimination of target speech extraction with time-domain speakerbeam,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.689464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.484727Z digest=sha256:c599d488c414e026fd0f5f5911b5c7f53a88d058a0f8c24ba2f6370b8f97a486

Observation 089f216d-62a3-4e31-8c87-6a76468e0dd6 · outbound

This paper cites X-TaSNet: Robust and accu- rate time-domain speaker extraction network,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction X-TaSNet: Robust and accu- rate time-domain speaker extraction network,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.587802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.611493Z digest=sha256:a03523f9ab9f331078eb7fc6e1aaa6756153f9cc52dfc08d4c173de4002b11bd

Observation 589c60fe-fc3c-4f76-9965-eb81eda70371 · outbound

This paper cites SpEx: Multi-scale time domain speaker extraction network,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction SpEx: Multi-scale time domain speaker extraction network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.505337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.711218Z digest=sha256:ce26a00ccd795e9340e0af803faa4a79064a8bddb8f72c876e3247616cdcb766

Observation 76335c91-0715-4543-8483-70d8986113a5 · outbound

This paper cites Time domain audio visual speech separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Time domain audio visual speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.415813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.808411Z digest=sha256:351b6cb5a0b7d473b0e35918f963af346cb92a2aefa1cc471cbac075ba666286

Observation 0c140bb3-f781-4e02-b8e1-db34f6dc92da · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Muse: Multi-modal target speaker extraction with visual cues,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.306690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:29.877370Z digest=sha256:beb0381bf2e7be4167656111df01bcd1a638e150c41c0b20ae9669f58ae91bda

Observation e867124a-969b-445a-bd4d-3759b391df81 · outbound

This paper cites A universally- deployable ASR frontend for joint acoustic echo cancellation, speech enhancement, and voice separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction A universally- deployable ASR frontend for joint acoustic echo cancellation, speech enhancement, and voice separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.106973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.773923Z digest=sha256:b90c5c87366acce3270af3c3594e3fff46c116d0b863b86197646325c5f13fd2

Observation 37e0dac8-7426-44fa-b2d8-c9ddd98fec6c · outbound

This paper cites Multimodal SpeakerBeam: Single channel tar- get speech extraction with audio-visual speaker clues,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multimodal SpeakerBeam: Single channel tar- get speech extraction with audio-visual speaker clues,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.129372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.061113Z digest=sha256:40a63533ce0a9eae15755687fecd65bb656483c1b1bb76c168e6f7fc7655c38c

Observation 2e663395-b6bf-4442-822a-a60e879dae13 · outbound

This paper cites My lips are con- cealed: Audio-visual speech enhancement through obstruc- tions,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction My lips are con- cealed: Audio-visual speech enhancement through obstruc- tions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.023699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.138599Z digest=sha256:496bf15a04bc4d6c6fdaf7e9224a8aef0d38703b7fca29316e76bc991683ef3b

Observation cb1c1913-d5b7-431d-93fe-2c193154737a · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multimodal attention fusion for target speaker extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.895207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.204098Z digest=sha256:53ede5c03c09ff5278fbd64835e2b661f104cc72517a8d7d8b6689c51f51a588

Observation a5b151bd-736d-4985-b74a-7bb033376187 · outbound

This paper cites An overview of deep-learning-based audio- visual speech enhancement and separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction An overview of deep-learning-based audio- visual speech enhancement and separation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.752065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.339631Z digest=sha256:3ce1add347af9d22682ae2ab0b8e37879aa3b0bd766ab4b579dc937ba94582ce

Observation 6724d629-e7b5-4f52-b762-c08f87cf91be · outbound

This paper cites Moddrop: Adaptive multi-modal gesture recognition,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Moddrop: Adaptive multi-modal gesture recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.555767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.460601Z digest=sha256:74d14a72e060cb6f046034baccaee22c5afed93ed04be2d9ea531337a2d43d2e

Observation 8af16198-4809-4e05-88a0-6afd2e435f73 · outbound

This paper cites Modality dropout for im- proved performance-driven talking faces,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Modality dropout for im- proved performance-driven talking faces,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.395794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.590018Z digest=sha256:ab7e5c3f6e9fadcd74e521f730467b42dcd95047b8ba0983838effab961bb0ea

Observation 7abef295-032d-4e16-965f-2dc171142a5c · outbound

This paper cites Learnable irrele- vant modality dropout for multimodal action recognition on modality-specific annotated videos,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Learnable irrele- vant modality dropout for multimodal action recognition on modality-specific annotated videos,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.255162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.693010Z digest=sha256:fba07f1d1eb2a54bfb476db27ee04cf3331abaff919bc263ed7fdcbb7151c8da

Observation 4a63c721-16c5-478e-a033-38feec9b9beb · outbound

This paper cites New Insights on Target Speaker Extraction.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction New Insights on Target Speaker Extraction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:31.507782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:31.507782Z digest=sha256:6cb6e3a87467e93a58fc98f3050d1a5d7de403de2355fb266b03d606cc7cb3b0

Observation c3ebc68b-ad67-43a7-90c9-503d5657ccb0 · outbound

This paper cites Multi- stage speaker extraction with utterance and frame-level refer- ence signals,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multi- stage speaker extraction with utterance and frame-level refer- ence signals,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.975622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.858657Z digest=sha256:14b1eb8f17838d5fb1b52d63ded50eccd9394aed8d4dc1b4627d94308f34c89a

Observation facb3b7c-3bab-4864-87d1-b872e9bf72b3 · outbound

This paper cites TasNet: Time-domain audio sep- aration network for real-time, single-channel speech separa- tion,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction TasNet: Time-domain audio sep- aration network for real-time, single-channel speech separa- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.850217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:30.988388Z digest=sha256:ecf46c1e94de1834970f1574ee15f35f4b5a0db9c811b4012414580d80116ee5

Observation 72479576-39ad-46d6-89b4-d5b5cc3a851e · outbound

This paper cites SDR– half-baked or well done?.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction SDR– half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.728864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.095766Z digest=sha256:14c3a67d68223ff6680260f2dd1d68564ea123acce573213ef140c8c6951c71d

Observation 09c07bce-6864-4834-8d42-672dee8ab427 · outbound

This paper cites Dual-path rnn: Effi- cient long sequence modeling for time-domain single-channel speech separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Dual-path rnn: Effi- cient long sequence modeling for time-domain single-channel speech separation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.622331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.187299Z digest=sha256:5966ecafaeeaabc673055164ebb55c4549e0a0f6155176a72c5108215f991190

Observation b30c94ea-a70a-4bd7-b763-553e6fd65efe · outbound

This paper cites Combining residual net- works with lstms for lipreading,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Combining residual net- works with lstms for lipreading,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.537888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.222323Z digest=sha256:dbd8bfb4e7eab2999f5c83082f87c4fb15e89d9ca4f36a5f40243428b682d3f7

Observation 343025d0-6147-415e-81b8-2125e2ea9722 · outbound

This paper cites Deep residual learning for image recognition,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Deep residual learning for image recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.424559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.278099Z digest=sha256:5609498a52aeee0203e66c784a49cf5d2beadc07557bbe7658cc69b756127b20

Observation b9272abf-7e09-4bae-990e-c5f33dc78deb · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction LRS3-TED: a large-scale dataset for visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:31.413980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:31.413980Z digest=sha256:0083d8e2cc4ce55bcf517ddaa04599e546dcccc3a3a10bc648f381fc7816c368

Observation 8a47bfce-57e9-4d7a-b72c-c28ddb18c00b · outbound

This paper cites Adam: A method for stochastic op- timization,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Adam: A method for stochastic op- timization,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.303545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.567625Z digest=sha256:28834ca2330316858cb54013d2d0b623ce5e5331d87eaea79a43535dda9858e4

Observation 5dada267-3769-4780-852e-2c6a11c20b64 · outbound

This paper cites Wavesplit: End-to-end speech separation by speaker clustering,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Wavesplit: End-to-end speech separation by speaker clustering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.159095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.614386Z digest=sha256:b01139f4603d0898cb4e564aa16e515c37a9fc4ddfc81777e97eb59418f206d1

Observation 003b5539-8b65-403b-9361-20b4f280dca3 · outbound

This paper cites Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.028738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:31.657945Z digest=sha256:0651added8de0daf5f0b737a352c696eee2755fd7f2bb63c7e4cc8ac4d3a0676

Observation c2afe8af-eead-4367-8380-18eb5a4428ea · outbound

This paper cites Layer Normalization.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Layer Normalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:31.716604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:31.716604Z digest=sha256:7e987966a169858dc91dfa00ef8fe1748f0b7a7b11dcf4320556a0309f9eea55

Observation f91be2b0-5ace-46f4-8264-ffdeaf227828 · outbound

This paper cites The inter- and intra-chunk RNNs are realized in the non-causal configuration using bi-directional long short-term mem- ory (LSTM).

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The inter- and intra-chunk RNNs are realized in the non-causal configuration using bi-directional long short-term mem- ory (LSTM)

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:35.327944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.949405Z digest=sha256:71170fb9136e73aef153c7f97081a935bdb4d314d74ce0ce11873004349c6b0e

Pith citing papers

Observation f5f5636b-395b-42d6-9ec2-ed1836a61fe9 · inbound

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction cites this paper.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:04:31.889166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:04:28.526024Z digest=sha256:16712a8831550e572e190490778108d09e930e6d4d6a17cacf5067975c0ce8b9