Pith. sign in

Paper Citation Record · LEDGER

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.00466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00466 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.836014Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:41:16.865462Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:41:18.737055Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bc195b6e-d5d0-4f63-a9d0-04ac866c9326 · outbound

This paper cites Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narra- tive speech.Current Biology, 28(5):803–809,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narra- tive speech.Current Biology, 28(5):803–809,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.666999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.230964Z digest=sha256:b0a12c4fc7d0c5470698a042d554bfd800013a84b43a4fce275b0296bd4ce6bb

Observation 8f307be1-86c4-49ac-8956-6f9759ca9460 · outbound

This paper cites Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.114926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.691413Z digest=sha256:4326fd0832c852962dbe8df897966b5588198b8dee58cca827bf5c6c0fca0a86

Observation 13523b52-39f6-431a-b171-19ca9d351e34 · outbound

This paper cites L-spex: Localized target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction L-spex: Localized target speaker extraction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.756034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.922931Z digest=sha256:6f765764f0690ad4422939ef0495328e5c187c05aef66df4150fabd6618b24a9

Observation e2974a40-42d1-4de0-8110-e608a5038a5f · outbound

This paper cites The cocktail party problem.Neural computation, 17(9):1875– 1902,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction The cocktail party problem.Neural computation, 17(9):1875– 1902,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.856502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.321264Z digest=sha256:3edfd28a7ffd60202a21d6ec797d7ccd58cf496dd80da5e7505818cccbe6aa80

Observation df764bc2-ffa9-4846-8928-5a6b8931a09f · outbound

This paper cites Speaker-independent brain enhanced speech denoising.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speaker-independent brain enhanced speech denoising

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.420567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.542157Z digest=sha256:14b88ed0df28b05494ccc34a7742fb85dde2f4d6677fded9968a32426da8d4d6

Observation 3c2dcafd-2de0-422b-af07-0b1cda3acfb3 · outbound

This paper cites Cross-modal global interaction and local alignment for audio-visual speech recognition.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Cross-modal global interaction and local alignment for audio-visual speech recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.007943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.738422Z digest=sha256:d5fa5a00d1a21d78555621537b52b5a0282aef30e3625881dcb6d00cd11deef9

Observation 5fa6fbf7-6719-4168-b5c3-cc96e9102fa3 · outbound

This paper cites [Le Rouxet al., 2019 ] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction [Le Rouxet al., 2019 ] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.729131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.828302Z digest=sha256:632f966f07c74783f9d1d0f621118c1653d251d7bac7b0166a7ace39508a1a1d

Observation efeb095b-13d3-4666-b5b3-3425874acfd2 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.910755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.910755Z digest=sha256:6c0930ec2b3e2b9776d51d5ef7962493e8371478ebfc20ee6323464ee13958c5

Observation 2efd5f25-00e7-4d3f-8874-eb720ed719a3 · outbound

This paper cites Audio-visual active speaker extraction for sparsely overlapped multi-talker speech.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Audio-visual active speaker extraction for sparsely overlapped multi-talker speech

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.476889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.002599Z digest=sha256:9f85bc444ed75450a7e56315fa57bc99041b73a40e7d11b917f3945703a4da53

Observation 51088153-38b9-44cf-9d43-d86685d3973a · outbound

This paper cites Av- sepformer: Cross-attention sepformer for audio-visual tar- get speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Av- sepformer: Cross-attention sepformer for audio-visual tar- get speaker extraction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.207261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.099123Z digest=sha256:69caa3d5bb032364dfc472137dd78654bffdb726a1fd991530eda18333930677

Observation bafb6e57-f5e0-45ae-8f6c-cb2825084cc9 · outbound

This paper cites Development of the audi- tory system.Handbook of clinical neurology, 129:55–72,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Development of the audi- tory system.Handbook of clinical neurology, 129:55–72,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.995877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.177957Z digest=sha256:18af26cd387b66b95203b0e9273012b910349e38bd19762730084bacf9d68222

Observation 5bd4c207-95ca-4561-bd42-5119b2fd12e1 · outbound

This paper cites Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.631601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.365027Z digest=sha256:497ec6667fd6dea6216230d7cc124bd505d22b2580ccc39add51f96acb6baf7d

Observation 1556424d-5a68-4f54-a3d0-172c329858fa · outbound

This paper cites Dbpnet: Dual- branch parallel network with temporal-frequency fusion for auditory attention detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Dbpnet: Dual- branch parallel network with temporal-frequency fusion for auditory attention detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.412865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.461928Z digest=sha256:6a4b789bdf74b41abd1ad8b5cd8e3dbd1b65c6ffb22efeacf2f33722b7a2ecd0

Observation 37e4535f-4d26-4e5c-9eeb-a1b405a74090 · outbound

This paper cites Attentional selection in a cocktail party environment can be decoded from single- trial eeg.Cerebral cortex, 25(7):1697–1706,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Attentional selection in a cocktail party environment can be decoded from single- trial eeg.Cerebral cortex, 25(7):1697–1706,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.240052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.586298Z digest=sha256:2cb57c70b76cbc701d76d8b606ea16f7b2122e047fe2af3ca54858f3febf20f3

Observation 6046fe9f-d1f6-4977-8a76-b8d3c9047317 · outbound

This paper cites Neural decoding of at- tentional selection in multi-speaker environments without access to clean sources.Journal of neural engineering, 14(5):056001,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neural decoding of at- tentional selection in multi-speaker environments without access to clean sources.Journal of neural engineering, 14(5):056001,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.058858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.677988Z digest=sha256:165e614e3b571bea94bf69277c1834cb07791c9dbd9bc43913c439255c082aef

Observation b3e6997e-0ba2-4970-ac42-73566d32e77a · outbound

This paper cites Neu- roheed+: Improving neuro-steered speaker extraction with joint auditory attention detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neu- roheed+: Improving neuro-steered speaker extraction with joint auditory attention detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.870687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.758386Z digest=sha256:1875c1cb18211001e3056dd63f24206e07129c4793556283005729a4d70cbd9d

Observation 20c99bcf-0126-4174-a528-1a75a08b5db6 · outbound

This paper cites Tf-nsse: A time–frequency domain neuro-steered speaker extractor.Applied Acous- tics, 211:109519,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Tf-nsse: A time–frequency domain neuro-steered speaker extractor.Applied Acous- tics, 211:109519,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.706750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.857130Z digest=sha256:b4156853d14a1f0a0918f20b60e4559ecfb9f347509e6e0f2a2f8f1880689145

Observation 6acfd61d-a7aa-40dc-a6c3-b881a3eb6a0e · outbound

This paper cites Phase space graph convolutional network for chaotic time series learning.IEEE Transactions on Industrial Informat- ics,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Phase space graph convolutional network for chaotic time series learning.IEEE Transactions on Industrial Informat- ics,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.492949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.950800Z digest=sha256:343dbdb93dfe0f6458ecb3935117e00f169c71e97bc855934b48ba9847ec900d

Observation 4beca81b-dc3b-475c-aa27-943dd8141bc5 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.286102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.035539Z digest=sha256:3b64d5721019dce6d7c238864ce9f9a338e0916666e684b2b5b8619a3f27413a

Observation 8e2a4f39-bf0f-4ec0-ae23-f64f8f361103 · outbound

This paper cites An algorithm for intelligibil- ity prediction of time–frequency weighted noisy speech.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction An algorithm for intelligibil- ity prediction of time–frequency weighted noisy speech

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.136099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.277497Z digest=sha256:6efce8e3530f8cf15ac1a39557c509372b7ba8a79226120104d8e5b4b7ba0e3a

Observation 3cb4f928-39ea-4f02-acbd-1f4a5fbd31e0 · outbound

This paper cites A study of multichannel spatiotemporal features and knowledge distillation on robust target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction A study of multichannel spatiotemporal features and knowledge distillation on robust target speaker extraction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.950548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.447655Z digest=sha256:cc0518d4a3a457b06aaafe00862c9d29e673cf9893c71247fd3ff7454be7f4ed

Observation ff249c62-b548-44b8-aa3d-a2595cb51d0a · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network.IEEE/ACM transactions on audio, speech, and language processing, 28:1370–1384,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Spex: Multi-scale time domain speaker extraction network.IEEE/ACM transactions on audio, speech, and language processing, 28:1370–1384,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.820826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.503332Z digest=sha256:65c3a4840a0cc1330c3774efa4aa282adfce9d7adec5387ea31abb188b5a7380

Observation 3ad76a10-e7c1-4e4d-8d48-72455b9e78c3 · outbound

This paper cites DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:17.955113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.613996Z digest=sha256:ca412db8a4577fde271ed913bb43d2cafa3dea0c8ca7254c0858fd91b2d5d26e

Observation bc8e5ec2-d71f-41e9-8950-eb2003d6de73 · outbound

This paper cites Basen: Time-domain brain-assisted speech enhancement network with convolutional cross at- tention in multi-talker conditions.Interspeech 2023,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Basen: Time-domain brain-assisted speech enhancement network with convolutional cross at- tention in multi-talker conditions.Interspeech 2023,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.705495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.692399Z digest=sha256:8391f4feaba685dd6bb8852e5e368be45fca480b3de929b1d4d1f151eb611986

Observation 4c96733a-00b6-4f10-8b93-6354c97f32d8 · outbound

This paper cites Based on audio-video evoked auditory attention detection electroencephalogram dataset.Journal of Tsinghua University (Science and Tech- nology), 64(11):1919–1926,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Based on audio-video evoked auditory attention detection electroencephalogram dataset.Journal of Tsinghua University (Science and Tech- nology), 64(11):1919–1926,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.568754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.749495Z digest=sha256:0dd0851d9028b599aec7a928cc64580b953f8b7dee5af9d63f9a77baea905f57

Observation 54d544ed-48a9-4f5b-b627-dfc15aff88a2 · outbound

This paper cites Neural target speech extraction: An overview.IEEE Signal Processing Magazine, 40(3):8–29, 2023.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neural target speech extraction: An overview.IEEE Signal Processing Magazine, 40(3):8–29, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.404454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.836014Z digest=sha256:54e7d6fa55e6196f75ac628270a3a105f3edf4669765ee93ea8d1a3eff965e12

Observation cb6d813a-1da5-4dd1-82c2-343b6f83e2c9 · outbound

This paper cites GroupMamba: Efficient Group-Based Visual State Space Model.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction GroupMamba: Efficient Group-Based Visual State Space Model

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.151270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.151270Z digest=sha256:fd3e1983a338c7aa24614206a45cd5063fe85199ffc8fc8a38400c3dfc77723f

Observation c4fc0c39-1540-4c36-83cd-f0c5369100b8 · outbound

This paper cites Centroid estimation with transformer-based speaker embedder for robust target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Centroid estimation with transformer-based speaker embedder for robust target speaker extraction

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.610268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.440078Z digest=sha256:c590a1678704542364d89be69827429a523ca86797caa8daaad6234a9018867d

Observation cc40db08-4309-4add-a24d-7547d36f37ff · outbound

This paper cites Speech intelligibility predicted from neural entrainment of the speech envelope.Journal of the Association for Re- search in Otolaryngology, 19:181–191,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speech intelligibility predicted from neural entrainment of the speech envelope.Journal of the Association for Re- search in Otolaryngology, 19:181–191,

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.047429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.369243Z digest=sha256:ee0d712c019ce29b6149179abe5aa619348906b0c2e7b4713d5b686ed97286bf

Observation 187771db-3a3e-4b65-a4cd-ec87fac5dad5 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256– 1266,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256– 1266,

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.818684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.284920Z digest=sha256:fe36eda7dad85c16780c0a2e9bf9a64a5bb09f73ed0111c2ed87d73e4c6995e6

Observation 493b7a6c-6905-47f2-88dd-dccd89182d9c · outbound

This paper cites Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception.NeuroImage, 223:117282,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception.NeuroImage, 223:117282,

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.313507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.335409Z digest=sha256:cbd5c533cc292a3a72780558488944ce8bc3b88f4b4e27d830e181901f9dccbd

Observation d0c08969-5790-4ea4-b3f3-29de1790de9c · outbound

This paper cites Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.107005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.107005Z digest=sha256:46743366edd6c8e33a4d5a0db9cc768c751210503723c0a6d92dbe87babcc206

Observation 0cd41983-8063-42d2-adc1-94a39ac04e1c · outbound

This paper cites NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.245325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.465323Z digest=sha256:931c81b0a7e3753240141601b16e452152495e7a94b9ebf3e9f8ba1f25053759

Observation f3f7cfbf-b1ee-46ac-8306-2586aa49a79c · outbound

This paper cites End-to-end brain-driven speech enhance- ment in multi-talker conditions.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1718– 1733,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction End-to-end brain-driven speech enhance- ment in multi-talker conditions.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1718– 1733,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.213036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.638340Z digest=sha256:a183fe70df1fd2bb2c7f553dcf04c0a26c7b18e30c5ad05a96031734f32f37aa

Observation 65930c91-b637-452e-9df8-32bf922a2c89 · outbound

This paper cites Speaker-independent auditory attention decoding with- out access to clean speech sources.Science advances, 5(5):eaav6134,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speaker-independent auditory attention decoding with- out access to clean speech sources.Science advances, 5(5):eaav6134,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.495225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.015545Z digest=sha256:66af4dc94f9212d958b1dc27aa629bf6738a97dd0c3843d98c85a68d5ab7d0b6

Observation f78d3635-634d-4020-9633-6f9d2e4ad98c · outbound

This paper cites X-tf-gridnet: A time–frequency domain target speaker extraction network with adaptive speaker embed- ding fusion.Information Fusion, 112:102550,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction X-tf-gridnet: A time–frequency domain target speaker extraction network with adaptive speaker embed- ding fusion.Information Fusion, 112:102550,

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.157869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.228471Z digest=sha256:c711e6f3622c8cee93aac1d62c7efe22404c599df19106e2c2baf382e7dde111

Observation e8195b1e-72b8-4d7c-aa33-933a0806b837 · outbound

This paper cites Msfnet: Multi-scale fusion net- work for brain-controlled speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Msfnet: Multi-scale fusion net- work for brain-controlled speaker extraction

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.082962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.580915Z digest=sha256:a64e2276269e752d98af9f572f926914bd9df7e165d30254bedc3fd6a02d5e2c

Observation b3fb19f1-d6c2-4a3c-ad58-871fcf6c22aa · outbound

This paper cites SpEx+: A Complete Time Domain Speaker Extraction Network.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction SpEx+: A Complete Time Domain Speaker Extraction Network

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:14.833000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:14.833000Z digest=sha256:3604636646406360836bf4d5888f3cffa3c67379f87e37aa71037d5c228ce30b

Pith citing papers

Observation 8de67a4d-dd98-4d1f-9a7e-c8cfb58efa2a · inbound

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction cites this paper.

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:41:18.762933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:41:16.865462Z digest=sha256:1c6c6639c426f499d22c064da6aa689b60650c6a23828d1fd3fa9193fe5edd68

Observation 36a42793-abee-433f-8157-6c307e1077df · inbound

Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss cites this paper.

Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T07:24:44.110448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:24:44.110448Z digest=sha256:2bb1997bdd4d81b51adf2843b130583d449e05b734dd475f1345db1df96ed246