Pith. sign in

Paper Citation Record · LEDGER

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

As of 17 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.06566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06566 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:31.716604Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:04:28.526024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:04:31.829062Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3c7bbe7-8cd5-4329-848c-3a3f824ebf9c · outbound

This paper cites Humans use auxiliary information, such as spatial and visual cues as well as speaker familiarity, to selectively attend to auditory stimuli [1].

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Humans use auxiliary information, such as spatial and visual cues as well as speaker familiarity, to selectively attend to auditory stimuli [1]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:36.275780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.438780Z digest=sha256:9b5d5571773dfc7f4bc42cdab9d8f1b073e2aafaa5c8ad81e8bca8f6ed1ee741

Observation f5f5636b-395b-42d6-9ec2-ed1836a61fe9 · outbound

This paper cites Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:04:31.889166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.526024Z digest=sha256:414237742e34bfc46a7badf4d4c56171291301b9669f28918d77ef0eeacf191e

Observation 4b9e1ac3-5e6a-4565-b09a-05a92e79abbb · outbound

This paper cites The architecture comprises an AudioClueNet module, a VideoClueNet module, an embedding combination module, and an extraction network.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The architecture comprises an AudioClueNet module, a VideoClueNet module, an embedding combination module, and an extraction network

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:36.013691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.619912Z digest=sha256:9b07e9e40a467d99d720d5450eaed6c6f93d02a4426b5d005d4c3346f90f465b

Observation 2b0a58df-6125-4f35-b39a-2c240e557696 · outbound

This paper cites 3 was trained using three differ- ent training strategies to study their effect on the model’s robustness.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction 3 was trained using three differ- ent training strategies to study their effect on the model’s robustness

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:35.798746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.736738Z digest=sha256:732b1ec82dffb78e65e3c432590bd4d1431711f05033ad4577d91fbb204c4fe6

Observation d8f53ced-b569-420d-8012-97098fbd582c · outbound

This paper cites Model description The basic building block of MTSE system under test is the dual- path recurrent neural network (DPRNN) proposed in [20].

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Model description The basic building block of MTSE system under test is the dual- path recurrent neural network (DPRNN) proposed in [20]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:35.558471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.858366Z digest=sha256:47b4285b7ca7934a759ccc6c3661e17c4bccc1c546a31bdf80c2481b66625f44

Observation 43779bf2-5a02-4247-a2cf-197af482189d · outbound

This paper cites Our initial experimentation showed the models to be sensitive to the normalization layers used in the DNN architecture.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Our initial experimentation showed the models to be sensitive to the normalization layers used in the DNN architecture

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T19:04:35.119973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.080203Z digest=sha256:b343cd091b04e2d0a10ad42f895ea50df1e373868b96fda31d7c43774082c0e5

Observation d4496882-b3c2-4409-b3c2-4ff9b4db4cee · outbound

This paper cites an unresolved cited work.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:04:34.999737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.182452Z digest=sha256:7454009779416e7217cebdf29e283736f2be85d6f8dd7611f085cd85125a1931

Observation 9a11ddb9-7eac-47d7-89ed-e135ddd24952 · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Usev: Universal speaker extraction with visual cue,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.204021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.965218Z digest=sha256:019114101fa4bcd4363e87f3bd4fe1923719b700e20fbff737c030ca49dd979e

Observation eb7c28ff-8c55-4f15-85d5-d271ef7cfc85 · outbound

This paper cites The cocktail-party problem revisited: Early processing and selection of multi-talker speech.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The cocktail-party problem revisited: Early processing and selection of multi-talker speech

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.896402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.282529Z digest=sha256:c706c3672d8217b0d29c374381d47c7c7dd98a50765ff6aece0091226d58127b

Observation 69977751-296b-4938-9819-bd2c7630f7d7 · outbound

This paper cites Single channel target speaker extraction and recognition with speaker beam,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Single channel target speaker extraction and recognition with speaker beam,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.794652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.364193Z digest=sha256:7e68654bfe6a64ff9e12fe8cf7e45cd410c85cded68755b4f8f8f49e88035905

Observation b9d8668d-4ecf-4fa4-8eae-1bb253f96e93 · outbound

This paper cites Improving speaker discrimination of target speech extraction with time-domain speakerbeam,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Improving speaker discrimination of target speech extraction with time-domain speakerbeam,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.689464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.484727Z digest=sha256:3ae8fdad7fd8d12e53d86529b49e9cfbc5fe6bd4183363d28330c4335d8f0214

Observation 089f216d-62a3-4e31-8c87-6a76468e0dd6 · outbound

This paper cites X-TaSNet: Robust and accu- rate time-domain speaker extraction network,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction X-TaSNet: Robust and accu- rate time-domain speaker extraction network,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.587802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.611493Z digest=sha256:3aae71ced18b7373e1df130caf0165cc9f6d0c422c0ea45e73476a6fa5d34f33

Observation 589c60fe-fc3c-4f76-9965-eb81eda70371 · outbound

This paper cites SpEx: Multi-scale time domain speaker extraction network,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction SpEx: Multi-scale time domain speaker extraction network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.505337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.711218Z digest=sha256:01f94bf07722a4e3495b4b3d6d9070cd576d10f9b106d984530d3c7e7bf2541f

Observation 76335c91-0715-4543-8483-70d8986113a5 · outbound

This paper cites Time domain audio visual speech separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Time domain audio visual speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.415813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.808411Z digest=sha256:b4a9b5fb90ca495d190a6e0753f8d2254681fad2d9caad03d981cb6e1b1606b2

Observation 0c140bb3-f781-4e02-b8e1-db34f6dc92da · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Muse: Multi-modal target speaker extraction with visual cues,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.306690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:29.877370Z digest=sha256:dbd97e6c32786a0cc7d8d8cf46cac46f69eb95248edf683c8617bd6a8e2d0c00

Observation e867124a-969b-445a-bd4d-3759b391df81 · outbound

This paper cites A universally- deployable ASR frontend for joint acoustic echo cancellation, speech enhancement, and voice separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction A universally- deployable ASR frontend for joint acoustic echo cancellation, speech enhancement, and voice separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.106973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.773923Z digest=sha256:394254708407e4acd86b51865a20fa50cd7d0fdb9b553a4657565b337f6b89e5

Observation 37e0dac8-7426-44fa-b2d8-c9ddd98fec6c · outbound

This paper cites Multimodal SpeakerBeam: Single channel tar- get speech extraction with audio-visual speaker clues,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multimodal SpeakerBeam: Single channel tar- get speech extraction with audio-visual speaker clues,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.129372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.061113Z digest=sha256:564cec3cb3f3bb7ec7bb0f014cf4baaffc9d9d4fb431f1b8b1b1c9e8caea6c2c

Observation 2e663395-b6bf-4442-822a-a60e879dae13 · outbound

This paper cites My lips are con- cealed: Audio-visual speech enhancement through obstruc- tions,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction My lips are con- cealed: Audio-visual speech enhancement through obstruc- tions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:34.023699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.138599Z digest=sha256:849d18a48d7e2763940887c76750cc5750bf04c299a9e269b7033c83d85bad5b

Observation cb1c1913-d5b7-431d-93fe-2c193154737a · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multimodal attention fusion for target speaker extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.895207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.204098Z digest=sha256:cc77bb15966e9477fa1473e9f154019fe7e961ab8a3f6f55cc29d60ee6411770

Observation a5b151bd-736d-4985-b74a-7bb033376187 · outbound

This paper cites An overview of deep-learning-based audio- visual speech enhancement and separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction An overview of deep-learning-based audio- visual speech enhancement and separation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.752065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.339631Z digest=sha256:d5b5ee4bb3e4590d3f6ff20a9c13e42ef5cfe74357fcd0a6ecd7dba4976223c4

Observation 6724d629-e7b5-4f52-b762-c08f87cf91be · outbound

This paper cites Moddrop: Adaptive multi-modal gesture recognition,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Moddrop: Adaptive multi-modal gesture recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.555767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.460601Z digest=sha256:8d635414911ba611fb60eaffc8f917713bdc93a5ef6c64651e7d11541bdc3fd6

Observation 8af16198-4809-4e05-88a0-6afd2e435f73 · outbound

This paper cites Modality dropout for im- proved performance-driven talking faces,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Modality dropout for im- proved performance-driven talking faces,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.395794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.590018Z digest=sha256:8ebaf001af40fd22435d729bd58d256c1baf6fafa18c9b8f1d633eccf396b1c6

Observation 7abef295-032d-4e16-965f-2dc171142a5c · outbound

This paper cites Learnable irrele- vant modality dropout for multimodal action recognition on modality-specific annotated videos,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Learnable irrele- vant modality dropout for multimodal action recognition on modality-specific annotated videos,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:33.255162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.693010Z digest=sha256:b711f1dbbc03833df9e19132d1592a5245887889bf3cfd5563f8d4126b2042ae

Observation 4a63c721-16c5-478e-a033-38feec9b9beb · outbound

This paper cites New Insights on Target Speaker Extraction.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction New Insights on Target Speaker Extraction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:31.507782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:31.507782Z digest=sha256:6cb6e3a87467e93a58fc98f3050d1a5d7de403de2355fb266b03d606cc7cb3b0

Observation c3ebc68b-ad67-43a7-90c9-503d5657ccb0 · outbound

This paper cites Multi- stage speaker extraction with utterance and frame-level refer- ence signals,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Multi- stage speaker extraction with utterance and frame-level refer- ence signals,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.975622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.858657Z digest=sha256:da4d82b9f2be1966a8f931fe9a5c00698119ad8e9bf9ccf1fedbaf2212eaa5e8

Observation facb3b7c-3bab-4864-87d1-b872e9bf72b3 · outbound

This paper cites TasNet: Time-domain audio sep- aration network for real-time, single-channel speech separa- tion,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction TasNet: Time-domain audio sep- aration network for real-time, single-channel speech separa- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.850217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:30.988388Z digest=sha256:b05adcb1740c65e9edba20165f0e44dd9758d1ed3743dbb79e07451baeec2818

Observation 72479576-39ad-46d6-89b4-d5b5cc3a851e · outbound

This paper cites SDR– half-baked or well done?.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction SDR– half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.728864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.095766Z digest=sha256:ae1072886a00669c86ef30e176485cf16b8c40e4f161bf829d3f82d121a194dd

Observation 09c07bce-6864-4834-8d42-672dee8ab427 · outbound

This paper cites Dual-path rnn: Effi- cient long sequence modeling for time-domain single-channel speech separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Dual-path rnn: Effi- cient long sequence modeling for time-domain single-channel speech separation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.622331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.187299Z digest=sha256:ff4065b07c50ce4c7aeb71f12175bd7808871b4ad1c5faf56245029c54b447ed

Observation b30c94ea-a70a-4bd7-b763-553e6fd65efe · outbound

This paper cites Combining residual net- works with lstms for lipreading,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Combining residual net- works with lstms for lipreading,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.537888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.222323Z digest=sha256:c6637848a988bb2a8e24528d6c811a0bbd2cbb85e002c79d8d0709360aca2436

Observation 343025d0-6147-415e-81b8-2125e2ea9722 · outbound

This paper cites Deep residual learning for image recognition,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Deep residual learning for image recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.424559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.278099Z digest=sha256:8ddf602e6335554546fe96a3d19900d2ab603429baa20477717c98b646064053

Observation b9272abf-7e09-4bae-990e-c5f33dc78deb · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction LRS3-TED: a large-scale dataset for visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:31.413980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:31.413980Z digest=sha256:0083d8e2cc4ce55bcf517ddaa04599e546dcccc3a3a10bc648f381fc7816c368

Observation 8a47bfce-57e9-4d7a-b72c-c28ddb18c00b · outbound

This paper cites Adam: A method for stochastic op- timization,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Adam: A method for stochastic op- timization,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.303545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.567625Z digest=sha256:07e685de6d6ebcabee2a1851c11f29d84891565434fc530bc6a9bf136318af64

Observation 5dada267-3769-4780-852e-2c6a11c20b64 · outbound

This paper cites Wavesplit: End-to-end speech separation by speaker clustering,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Wavesplit: End-to-end speech separation by speaker clustering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.159095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.614386Z digest=sha256:70a61ce149fac7de7a944dbe6f879cf87404992c912e882b13447ebb9ed79b7c

Observation 003b5539-8b65-403b-9361-20b4f280dca3 · outbound

This paper cites Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:32.028738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:31.657945Z digest=sha256:17de2e9195a213b60a0e11cc229193bb975f3394e630c406bf736b5757b0a142

Observation c2afe8af-eead-4367-8380-18eb5a4428ea · outbound

This paper cites Layer Normalization.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Layer Normalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:31.716604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:31.716604Z digest=sha256:7e987966a169858dc91dfa00ef8fe1748f0b7a7b11dcf4320556a0309f9eea55

Observation f91be2b0-5ace-46f4-8264-ffdeaf227828 · outbound

This paper cites The inter- and intra-chunk RNNs are realized in the non-causal configuration using bi-directional long short-term mem- ory (LSTM).

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction The inter- and intra-chunk RNNs are realized in the non-causal configuration using bi-directional long short-term mem- ory (LSTM)

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:04:35.327944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.949405Z digest=sha256:b5032c496fd6646921db1290531f27498efd97edad3e99c33db565ec8a545fe6

Pith citing papers

Observation f5f5636b-395b-42d6-9ec2-ed1836a61fe9 · inbound

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction cites this paper.

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:04:31.889166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:04:28.526024Z digest=sha256:414237742e34bfc46a7badf4d4c56171291301b9669f28918d77ef0eeacf191e