Pith. sign in

Paper Citation Record · LEDGER

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

As of 17 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2507.15294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15294 v2

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:09.564250Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:39:50.277622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy72
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83c6922e-33a6-4d1b-9c21-65841d11fc41 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Some experiments on the recognition of speech, with one and with two ears,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.175918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.175918Z digest=sha256:9cfd669f98a959df803e858a4b72563c24a03f9dfde8170911d69d3a204007e0

Observation 744f2f48-ebd7-4a45-8006-d2f53eb17f8b · outbound

This paper cites The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.181286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.181286Z digest=sha256:04a2f8c415280fc8db16dc445e908bec7ba9aef2f988ed3a9796b7cda484227b

Observation 8b3c2740-4d0c-4fe9-9692-f331dc2aa051 · outbound

This paper cites Seeing to hear better: evidence for early audio-visual interactions in speech identification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Seeing to hear better: evidence for early audio-visual interactions in speech identification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.949415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.186675Z digest=sha256:2d9ef5a6d2ff7dbf31a9a2a598ecc0c6c3d216293fd06413d10dbad9c71a51a8

Observation 95ec1176-f5ff-4daf-aed3-5b0f40666070 · outbound

This paper cites Visual contribution to speech intelligibility in noise,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Visual contribution to speech intelligibility in noise,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.932674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.191605Z digest=sha256:f19585ee1fa34e08a706ff42cdd66c72317e5d4ded20539cfeb23d74a94d3800

Observation 4ff5073d-3358-4e53-8caa-ed50b21744a5 · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Muse: Multi-modal target speaker extraction with visual cues,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.915156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.196835Z digest=sha256:374042e8421b107ff8ab562dc4eb0ee95eece6c9f5672504cbcea4913b240478

Observation 2763e41b-075b-4632-925b-f5e4921a5a8c · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Usev: Universal speaker extraction with visual cue,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.885618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.201540Z digest=sha256:42ec6b044e691f371ed4fd608a4fcc912c0e3155679cc7ecd44fa8822641755b

Observation d1c517a4-72d4-4dbd-b3a0-a2c100695cab · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.866995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.207358Z digest=sha256:3405753a4bd3abf14343817dc70ca33c42853829a4d3d6d85bf0b44438502a58

Observation 094b854b-5973-4c6b-a35e-cf83b49c2596 · outbound

This paper cites Neural target speech extraction: An overview,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neural target speech extraction: An overview,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.212164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.212164Z digest=sha256:03b49662cd4343eb0293be531d2403cbf217680384fd92bc1d87c52615a37fb0

Observation a76edd7b-75fe-4c81-b09d-a2a30fc5197c · outbound

This paper cites An overview of deep-learning-based audio-visual speech en- hancement and separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An overview of deep-learning-based audio-visual speech en- hancement and separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.841856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.217781Z digest=sha256:2ba06115fb99f7ed95c0d2a2b7f3b1005ac559f05a3dc22ea97ccdab88470395

Observation 1a4b2d8f-8080-4e8e-a875-1c91d982f1bd · outbound

This paper cites PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.826667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.222557Z digest=sha256:691a31ef24c8ed2292339e78c8cdf8def78b6ac6355e0763b0f05d6f11aad95b

Observation 7d13eadd-9117-4305-83c8-634b94bccd40 · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speaker extraction with co-speech gestures cue,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.809836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.227254Z digest=sha256:c568e6e31c0f51506e3415a3bbcbd355e2b0ea8b66e2bdcd119f26f1bf027fd8

Observation 0f02fc49-23af-4400-a7a3-c539854bd177 · outbound

This paper cites Rethinking the Visual Cues in Audio-Visual Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Rethinking the Visual Cues in Audio-Visual Speaker Extraction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.792840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.231505Z digest=sha256:315cb5f9aa1aff87721816511d5b3e5c8bf92340f75e34941f16fd14ffb86ade

Observation 302839d1-bece-43a4-8150-2868d3c52695 · outbound

This paper cites FaceFilter: Audio- Visual Speech Separation Using Still Images,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions FaceFilter: Audio- Visual Speech Separation Using Still Images,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.767074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.235980Z digest=sha256:e8e5609777b8b28eb2524c118453133d54c90b66126c1148a43033a38cc47f50

Observation 91365b9e-2036-43e3-b47e-3086683cfa8c · outbound

This paper cites c 2av-tse: Context and confidence-aware audio visual target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions c 2av-tse: Context and confidence-aware audio visual target speaker extraction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.751607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.239987Z digest=sha256:d3d41638a63e1fca532f820e90c425ca518a94479e05c6b9e5ecfd07aa774d66

Observation 0a5ad71e-687a-4064-b397-262808e4d4ae · outbound

This paper cites Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.736790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.244151Z digest=sha256:b975a673d7f2004ab92f6eb14a165c402c37984c7975e58475026e752afa97b4

Observation 982b878b-e97d-4872-a5db-3910f7170acb · outbound

This paper cites Iianet: An intra- and inter-modality attention network for audio-visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iianet: An intra- and inter-modality attention network for audio-visual speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.721532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.248745Z digest=sha256:1200d77231c6d69b925c2202c0ea1dd4b825bdbacffc13a52ee5777e3386a301

Observation 68eefdee-0a10-4910-844e-7f8f666b21d1 · outbound

This paper cites Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.704696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.253262Z digest=sha256:ce4c44c716a1dc3b21bd9d5d5141bef55616e4930ad9b5ec3972a0448b1c7339

Observation 5516288b-28fe-4e07-9780-6485d57fdb4e · outbound

This paper cites Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.688983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.257996Z digest=sha256:e0ca807310b49f8d8e019de9d9d3aad0c3641a0150d729e3bb098e4fddaadc44

Observation 251799e1-7c01-4669-abea-64bcb8f411b9 · outbound

This paper cites Noise- disentanglement metric learning for robust speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Noise- disentanglement metric learning for robust speaker verification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.672219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.263291Z digest=sha256:25f5b79e533a3654fe51b66b363136b7a6968710b40c0a52afbd922578e4e859

Observation 9b109bc0-8f1d-4e45-9b0c-b6a4cc57e48f · outbound

This paper cites Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.656814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.267901Z digest=sha256:0cde01fc042d09c23cab5bc7761bb8ffc2ba3450f4f6ca95b7bf144092d2b805

Observation 0cb45b8b-12b0-4d2d-bc15-9b7e1d9cb579 · outbound

This paper cites Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.641107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.273119Z digest=sha256:6d4d33c0f82858a8c0e5ce865214dba80e0cf2112a7222f263ed5ad0e84e2ec3

Observation ca417b6c-9d8f-497e-817a-1d9c3321d122 · outbound

This paper cites Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.621867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.277199Z digest=sha256:2d555455f1a75b891c2620dea6dc1d1b8d8cc9d48ea35fad8feb3df94ebf65fd

Observation d4283792-3a6a-4479-af6c-1321dd6c1cc4 · outbound

This paper cites Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.606055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.281465Z digest=sha256:dcf787074c29b66596d1611773f956a156006295f7b39da505b2bfa1a3022a7a

Observation 5f16851f-04db-428a-988a-61d4a7da1a78 · outbound

This paper cites Time-domain audio-visual speech separation on low quality videos,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time-domain audio-visual speech separation on low quality videos,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.590626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.286027Z digest=sha256:a12a7752db078c2f7b11587933a06a0904ae90b7294b0752543c6576e08611f1

Observation 952af7c7-1dfe-4740-b23e-d0bd0ab8c3a9 · outbound

This paper cites Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.574988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.291033Z digest=sha256:9e1347c517c7521e50c96283e8e4ae8597e4b912092ca6ee2195df0a816ed453

Observation e7dcd13f-e53e-4227-a762-1dd98aa07395 · outbound

This paper cites My lips are concealed: Audio-visual speech enhancement through obstructions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions My lips are concealed: Audio-visual speech enhancement through obstructions,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.559734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.296351Z digest=sha256:7ec489fa190af862e1318890402abb76e8bb982d4c92910f8c517c645f26e831

Observation 6619f601-ea29-4e01-9e11-6d48e8324806 · outbound

This paper cites Multi-cue guided semi-supervised learning toward target speaker separation in real environments,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-cue guided semi-supervised learning toward target speaker separation in real environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.545050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.300822Z digest=sha256:d678ecc8cf35f143fb67b1cb83191a653fb48378455e5305c7d3d82e5327c63c

Observation f4f9bccf-176a-4260-a95b-6ca2b83a4649 · outbound

This paper cites A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.528996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.305414Z digest=sha256:0d4713b0718f26a6c6031d6a95aeee654e9782660b2a257ffebc49b5d9ce05c5

Observation 84bce1f1-23d6-4331-9ed0-991756082d7f · outbound

This paper cites Cross-modal speech separation without visual information during testing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cross-modal speech separation without visual information during testing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.513641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.311071Z digest=sha256:3670ee747da5411bee64c904e0122037aea4de8c202748fdeb1bd00053c51718

Observation d62436f8-6c39-4acd-acbb-d67d25f39ccc · outbound

This paper cites Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.498357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.317501Z digest=sha256:a8986cadc83cda74240ddc206d3027135430ed3f8266e6a5703beb252231ef05

Observation c459807a-0a79-4ed6-8058-0cea66615d59 · outbound

This paper cites Cocktail party listening in a dynamic multitalker environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cocktail party listening in a dynamic multitalker environment,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.482734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.323438Z digest=sha256:351f0d8f1f389967f8ae873a2678f9bca4ea2f66fe84a45852285207298d64b3

Observation e408aa7c-a002-478f-a24e-080c71069d95 · outbound

This paper cites The advantage of knowing where to listen,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The advantage of knowing where to listen,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.467314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.329382Z digest=sha256:e62854d5cc1c0f217f25ed6b661cf4114e86dd8c24aff1dd64aa5fafcaedde27

Observation c7f82ef1-b792-452b-a6db-e60f3923d14c · outbound

This paper cites Object continuity enhances selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Object continuity enhances selective auditory attention,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.450245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.334812Z digest=sha256:755f3bd91f4cbdce43bea25c578944aaaf41357753bab862f4757715f6c9f557

Observation 67e2d8c8-66db-4084-b3cc-1a7574113c4e · outbound

This paper cites Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.432061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.339775Z digest=sha256:7ffed74dc013087c4ffda01edbb95f7cb6be1c70539c6cd7fdcf9a638b4d16b1

Observation c202bef3-25b1-4d8c-a92a-f2e3d8b3cbbe · outbound

This paper cites The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.414053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.344866Z digest=sha256:2bc65909d29595273c122740ceb27d30f9d162bcff7da9e6d10f7f1de8f50518

Observation 92d5150e-1fa1-4fa2-8fe1-bc72d4bc2acf · outbound

This paper cites Attention and working memory in human auditory cortex,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention and working memory in human auditory cortex,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.395278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.350053Z digest=sha256:1c6deebceabf023ab5078521dacda661a3d61470184a7c5d41dffdc14ecb2a6d

Observation 877109a9-6ee5-418f-aefb-59121d7b42b2 · outbound

This paper cites Interactions between attention and working memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Interactions between attention and working memory,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.377778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.354448Z digest=sha256:3794b4ba77f668ef044a816072c039fe15b2bf9221a68a208b22d7bf40c3f79a

Observation f1508676-d493-4fde-b68c-adfcae7d7c55 · outbound

This paper cites Look once to hear: Target speech hearing with noisy examples,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Look once to hear: Target speech hearing with noisy examples,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.361411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.358543Z digest=sha256:927b736a602d10f36c183e930d09ef83b9cece75543e51f78164bfc18e68de12

Observation e077f5c0-d99c-48aa-9f92-fb7421d248c6 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal attention fusion for target speaker extraction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.362651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.362651Z digest=sha256:fa3042668767348000409b792d72de156ddfbc089dfdda4cc454f7b8969374b7

Observation d6d0b610-dd7f-440e-bebd-a6e46823b311 · outbound

This paper cites Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.332485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.366891Z digest=sha256:b23fa7d6d74f2b683e9a38db5774f0b4986a63691c2059f7ce4bba4c82cecc0f

Observation 5987cf6c-44af-40fa-a653-65ff48ffcdec · outbound

This paper cites Memory networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Memory networks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.311379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.371554Z digest=sha256:78f106a0a84e65239b272d2873c366968f3dfa67ad9c5681aabf095e7ef7915c

Observation 3b1f9cce-ca92-414a-97b1-3694b4dd8106 · outbound

This paper cites End-to-end memory net- works,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions End-to-end memory net- works,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.289641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.377652Z digest=sha256:e7308643b24c29a8a71d84c040ce93f88647692754ba2d433a62bf300fc485ef

Observation c9af338d-836f-4736-ac78-904d87b15541 · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Unsupervised feature learning via non-parametric instance discrimination,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.268598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.382343Z digest=sha256:37bbc92934639a64f50de529b7caf941198cf8c7a39b24a798ce32cc643941a1

Observation 91992967-d5e3-4a92-b2e2-d45a6a180914 · outbound

This paper cites Cromm-vsr: Cross-modal memory augmented visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cromm-vsr: Cross-modal memory augmented visual speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.250036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.386674Z digest=sha256:2ca713a64959fab9753e785664631d88bd419c06e3656d14b1d6f698d5b2f368

Observation f400ee2c-ab6b-4fc8-8a4a-bb3348d44d10 · outbound

This paper cites Multi-temporal lip-audio memory for visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-temporal lip-audio memory for visual speech recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.235219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.391493Z digest=sha256:17885df0f778a8fb18f8f68bf30725c8a54d96bcf198e185c99d209be385ce29

Observation 3d307f7b-3734-4902-b509-f0f6e5b525b0 · outbound

This paper cites Multi-modality associative bridging through memory: Speech sound recollected from face video,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-modality associative bridging through memory: Speech sound recollected from face video,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.220095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.395575Z digest=sha256:0ca00b41290dc18a8d627b023e41f7d604aebbefd4ff58a8836058382cf2b550

Observation b7540409-2519-41ea-b221-7034dc95db8b · outbound

This paper cites Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.202608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.400260Z digest=sha256:61ca8345290b94469b348e9942f92fb7679038e66eec469a879bd5b4fa78e282

Observation 62dbf9dc-1e99-402c-a311-d97e8e29a2b3 · outbound

This paper cites Distinguishing homophenes using multi-head visual-audio memory for lip reading,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Distinguishing homophenes using multi-head visual-audio memory for lip reading,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.184305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.405373Z digest=sha256:c51c14f54c9fc8e8101b9c2b3de05ed1474bf7b6e2346c2f3873f2e7158d26c9

Observation 8b8a75df-6b28-4901-82c0-914c3a569f20 · outbound

This paper cites Speech reconstruction with reminiscent sound via visual voice memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speech reconstruction with reminiscent sound via visual voice memory,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.166533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.409884Z digest=sha256:c15b835aebee16283859a69aca9b56939215d0152abf3e95011a75e6713afa1b

Observation aec87788-c5ef-4974-a09f-11d8b2aa3a8e · outbound

This paper cites Modeling attention and memory for auditory selection in a cocktail party environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Modeling attention and memory for auditory selection in a cocktail party environment,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.151632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.414371Z digest=sha256:4c6391f3d23ea3917b21fa41789b2771f00c5c1b823c7a5eff4508490bc79208

Observation 4057b1ff-ab00-44c2-a304-2d26fb5eb04f · outbound

This paper cites Explicit-memory multiresolution adaptive framework for speech and music separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Explicit-memory multiresolution adaptive framework for speech and music separation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.135206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.419236Z digest=sha256:e7eb2976194111c05eddb4c56038aa0faeb32ebe78417d4b45ba0cdd863ed052

Observation f8dedcdf-0717-4d31-ace6-d546dc222950 · outbound

This paper cites On the effectiveness of enrollment speech augmentation for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On the effectiveness of enrollment speech augmentation for target speaker extraction,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.118919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.423757Z digest=sha256:717c6166767e0d4d73dcb64746b36d59653600c09a7b8aa7a62682f0cbce482f

Observation 09f89054-bda5-48b8-8c9f-f5e021deb3c4 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Selective listening by synchronizing speech with lips,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.102594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.428367Z digest=sha256:9d32bfd164a260234fd4f5be3acb95a907f2e5bf8f19eca14bf9c8043daabddd

Observation 02d088bf-9e49-468c-b71e-80c73883c938 · outbound

This paper cites Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.086812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.432582Z digest=sha256:24e0afc49d58ec75c4a059da7d817ebf4d1975bdcd44866a2ef9117679c00dfa

Observation 7752d5ec-a8d0-4ad4-b71a-94b4668935d2 · outbound

This paper cites Listening and group- ing: an online autoregressive approach for monaural speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Listening and group- ing: an online autoregressive approach for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.070125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.437077Z digest=sha256:9d3abd49f661e7387e19c54042f0a69ccafc6c3ed727c039a8c223f32ea77321

Observation ec8da356-41c7-4aa4-ad74-9509c544109d · outbound

This paper cites Source-aware context network for single-channel multi-speaker speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Source-aware context network for single-channel multi-speaker speech separation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.053802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.441118Z digest=sha256:6141f27ca3dea4fbd0de1c0b6a2c7d86220ff7d2ef12938e278c7968667687e2

Observation 7b63a849-45b7-4346-b50a-5dfc98bd0baa · outbound

This paper cites An online speaker-aware speech separation approach based on time-domain representation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An online speaker-aware speech separation approach based on time-domain representation,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.037326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.445488Z digest=sha256:f4b2371254629a469c1ec40ba2aeb680beee3bb733a45d9be4f9146625d67f12

Observation c75e5f9b-344c-48e1-a81c-cbb23b9a1946 · outbound

This paper cites Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.020906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.450282Z digest=sha256:b6a319d98712c2a6df61b27bfd54064fd140b1fcf89a4d15c4304a9dcce9a20d

Observation f3f3203c-8165-4df5-8d2f-c873e09cbeab · outbound

This paper cites Neuroheed: Neuro- steered speaker extraction using eeg signals,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed: Neuro- steered speaker extraction using eeg signals,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.004328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.454651Z digest=sha256:8412ebd3dd49fff324c4b8e3b7dcf2743be38afca03be069ccf3a587a94c8a55

Observation e6b96c13-e4cd-4786-a305-f25608f9ce1e · outbound

This paper cites Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.985128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.458874Z digest=sha256:12a21dfcac51dc4e66482548f3461c097b408f9ff3b3939ec31769d906e60b76

Observation 0f2bcac3-be4f-490a-ab12-cecd3933a5e9 · outbound

This paper cites Paris: Pseudo-autoregressive siamese training for online speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Paris: Pseudo-autoregressive siamese training for online speech separation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.968849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.463998Z digest=sha256:af9ff8667b270b8a9a033ad14a6f3e898ae5531e17d67fcdf1e3b1e625777b32

Observation 900d7c57-838a-4484-a067-9c545a703533 · outbound

This paper cites On- line Audio-Visual Autoregressive Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On- line Audio-Visual Autoregressive Speaker Extraction,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.949809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.469257Z digest=sha256:8cbac3d6c07022af7459f91642a73ac1fb1430bec72530ea48c8b810dffab9e0

Observation d6aa6f40-20ba-4b58-94f3-1292d36b9aab · outbound

This paper cites Attention is all you need,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.473959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.473959Z digest=sha256:6f903b79d08589bba83611a36d6d71401b35db7bd79b3aeb983d6171322f64e3

Observation 81286334-2444-45bf-baeb-f40273a994ad · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.916620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.478541Z digest=sha256:45928e57950f938fa09ec0665c2ee47608ea90c58110c42528ce5144fb621b9d

Observation f360c19f-93a5-4926-8a80-b8ebe629387f · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.483332Z digest=sha256:8d89c911a4a02e74b390c1abca63b596a7aaf9ba38aeb5eb1a0ee5426099b3b5

Observation a5e18d5e-3975-41f5-83d8-e7782d961065 · outbound

This paper cites Curriculum learning,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Curriculum learning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.488014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.488014Z digest=sha256:01d969826ea13b2a2b61624a11c15c3c1eccdc70f04219ff816ceacb6b9b3ef8

Observation c59cc09d-4efc-4930-8960-00e534280cab · outbound

This paper cites A learning algorithm for continually running fully recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A learning algorithm for continually running fully recurrent neural networks,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.492392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.492392Z digest=sha256:b05400c1d5506f3410aa1a95b986562aac399fb0cf6839368ea2bc1c46991f1c

Observation 3b0191e3-e40d-42e9-88a2-1557279844d5 · outbound

This paper cites Scheduled sampling for sequence prediction with recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Scheduled sampling for sequence prediction with recurrent neural networks,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.858309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.497009Z digest=sha256:0366873fac8ae3cd2f75ce2a09d504dc58f6e62b6524ceed43dedf76d7443ea4

Observation 17078865-8ccd-4862-b091-1f5803386d60 · outbound

This paper cites Delving into high-quality syn- thetic face occlusion segmentation datasets,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Delving into high-quality syn- thetic face occlusion segmentation datasets,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.842492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.502291Z digest=sha256:51ac358ece58a7ebbf5a00af862e7a3da6b8e70bf5b0cfa061cf2d00903aa731

Observation ac2c226e-f775-4761-9c08-30a9a67f8a2a · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.825636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.506931Z digest=sha256:25e7d9a833fcb04e6b9caed724142f2a2c0661766b4ebe66bd99bbf0976bcd8f

Observation a38afe9a-a0f5-4d0f-b894-fb657b1d27c7 · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions V oxceleb2: Deep speaker recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.806766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.511270Z digest=sha256:2829b18822f37cc33f9dbadd3ab3001392909acafd4bdea9c0ed7612197f00ff

Observation 60b84ecc-6bd7-4797-a666-4cc5be932f54 · outbound

This paper cites Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.787987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.515654Z digest=sha256:1b4c16a543cd3fe1a7202f8dc9362c2eb91d2b0cf0d8aaff934d6cdd44f1a306

Observation 7e5d0481-f3bb-42df-8a8c-4fe34c912811 · outbound

This paper cites Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.766513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.521222Z digest=sha256:9dc1298892dfa6e71aca1249d905af5192b8cc538c162858d376b0b334f89e18

Observation 89aae3a2-1857-492b-bb66-7adf0a6cd7da · outbound

This paper cites Sdr–half-baked or well done?.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Sdr–half-baked or well done?

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.749611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.525852Z digest=sha256:75f2c9f272b2b99dc4524e11ba079d8b844f610a61e50e1d85068274a7659499

Observation 03e57a72-d5f5-46cf-8ff2-e8a5e8c6791b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.530574Z digest=sha256:508bc319630c11d6dea1c6a88aa2d83d28af0d0d4c4f9554a2b74eded31b3d7b

Observation c9560885-5aa4-4169-8ffe-8241c468c3dd · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.713313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.535064Z digest=sha256:25d3b62510862cb476b58320ccad9e9c0ab8d3740f9906af026b73cf76d30b05

Observation f1c2c8e9-7380-4dc0-aab9-0c7cabd69ed8 · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.693485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.540459Z digest=sha256:9adabc3152f2b74548c5537d07a94dc20cabc15339e7160659e405fc64504e81

Observation 7cf1d4a5-6276-456f-9ad5-bdb419330ec7 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.545087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.545087Z digest=sha256:5363676dd07537449a9ab33b36554a39d12323025a9aec4cf4ff2d25eb88075f

Observation 4453c919-dd08-4af9-a213-2da60ed02d89 · outbound

This paper cites Music source separation with band-split rnn,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Music source separation with band-split rnn,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.549514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.549514Z digest=sha256:9b41c6258caf04622c97d914b352de6028cf9fbd5c8db7952a4e7fbb71a28483

Observation 88f831c1-baf9-4a3f-9125-e6d48eb7bc4d · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.648914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.554592Z digest=sha256:3e98a271974d3eeb5d82f2e3783319ef8fa6d6b86e9be9235cdd19b1dd0a9af1

Observation 8e34c688-aeea-432b-b1fc-c5cc47cd47d5 · outbound

This paper cites Audio-visual target speaker extraction with selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Audio-visual target speaker extraction with selective auditory attention,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.631585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:40:09.559373Z digest=sha256:62c0f08a4985c4dc0ceb3f285733b1566a54d72204209f18aee97b11ee6b4ccb

Observation aa5cd0ee-cf22-4bee-8798-2d7dd572e36e · outbound

This paper cites Attention is all you need in speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need in speech separation,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.564250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.564250Z digest=sha256:d2454a9450c194e79b9d7216310ee80d0c1990bb7a5b501b07e91e8f3aa4eb15

Pith citing papers

Observation a093622f-23ed-4eee-b9d4-df82da779cc2 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.277622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.277622Z digest=sha256:8379d067dcbcee4943e353228cd97ea8c48c580b7466477cee871d330dcc0d3d