Pith. sign in

Paper Citation Record · LEDGER

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.00466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00466 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.836014Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:41:16.865462Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:41:18.737055Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bc195b6e-d5d0-4f63-a9d0-04ac866c9326 · outbound

This paper cites Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narra- tive speech.Current Biology, 28(5):803–809,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narra- tive speech.Current Biology, 28(5):803–809,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.666999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:14.230964Z digest=sha256:323f265b505e659bc41f1e3aee966224ad3bd354a1ffff33804667b83fb8704b

Observation 8f307be1-86c4-49ac-8956-6f9759ca9460 · outbound

This paper cites Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.114926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:14.691413Z digest=sha256:e5c061e4ab01dc8ad6902b1e3e7364941e6a9bd80cdc15838b0b8687e53d85b8

Observation 13523b52-39f6-431a-b171-19ca9d351e34 · outbound

This paper cites L-spex: Localized target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction L-spex: Localized target speaker extraction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.756034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:14.922931Z digest=sha256:0eb3a7f3f7f777676e50c9231e7e52a3cab30932f6c0754d771d060a2877b5d7

Observation e2974a40-42d1-4de0-8110-e608a5038a5f · outbound

This paper cites The cocktail party problem.Neural computation, 17(9):1875– 1902,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction The cocktail party problem.Neural computation, 17(9):1875– 1902,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.856502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.321264Z digest=sha256:062b0840b1675f8eedcc7647959135e6da4b197ee67652726c12b74ac132b606

Observation df764bc2-ffa9-4846-8928-5a6b8931a09f · outbound

This paper cites Speaker-independent brain enhanced speech denoising.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speaker-independent brain enhanced speech denoising

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.420567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.542157Z digest=sha256:dd65d110eedd8329810ae45cb7b247db3c3f8a66780f14a31268812c0c2505e4

Observation 3c2dcafd-2de0-422b-af07-0b1cda3acfb3 · outbound

This paper cites Cross-modal global interaction and local alignment for audio-visual speech recognition.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Cross-modal global interaction and local alignment for audio-visual speech recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.007943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.738422Z digest=sha256:02de2276180c4b857005cc66dc56015a6ecbc550e9c34237969fb09d6ee081f9

Observation 5fa6fbf7-6719-4168-b5c3-cc96e9102fa3 · outbound

This paper cites [Le Rouxet al., 2019 ] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction [Le Rouxet al., 2019 ] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.729131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.828302Z digest=sha256:ae151b2697b2eab98898be2cf74f590e639078e32d1bce5406ae289c5a4955e7

Observation efeb095b-13d3-4666-b5b3-3425874acfd2 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.910755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.910755Z digest=sha256:6a8a12076ae851ab775c223ff6545ca863f76386042a59fe5671885215916162

Observation 2efd5f25-00e7-4d3f-8874-eb720ed719a3 · outbound

This paper cites Audio-visual active speaker extraction for sparsely overlapped multi-talker speech.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Audio-visual active speaker extraction for sparsely overlapped multi-talker speech

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.476889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.002599Z digest=sha256:d7df659e7faabda33a05fde656a9b445f17ac4f9367444cb3043c7d2de3aa04d

Observation 51088153-38b9-44cf-9d43-d86685d3973a · outbound

This paper cites Av- sepformer: Cross-attention sepformer for audio-visual tar- get speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Av- sepformer: Cross-attention sepformer for audio-visual tar- get speaker extraction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.207261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.099123Z digest=sha256:cfcc4c82d9be7a48fe4490ebeda8926ac9a7d091c8d657bae45257c8ac7ca168

Observation bafb6e57-f5e0-45ae-8f6c-cb2825084cc9 · outbound

This paper cites Development of the audi- tory system.Handbook of clinical neurology, 129:55–72,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Development of the audi- tory system.Handbook of clinical neurology, 129:55–72,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.995877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.177957Z digest=sha256:6bd15a80686d98ffa387fe86af4f8ae437494a07b6b57eb60b68ae105eae8384

Observation 5bd4c207-95ca-4561-bd42-5119b2fd12e1 · outbound

This paper cites Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.631601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.365027Z digest=sha256:c805c7d76bba205440c20406e9ff014f4ee48592773d3f09fa03b55155fcb5a7

Observation 1556424d-5a68-4f54-a3d0-172c329858fa · outbound

This paper cites Dbpnet: Dual- branch parallel network with temporal-frequency fusion for auditory attention detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Dbpnet: Dual- branch parallel network with temporal-frequency fusion for auditory attention detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.412865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.461928Z digest=sha256:a55d1a7115b4803e239cb02a047f9af11fc8ffe6b4a9bad95c3adf65b9580130

Observation 37e4535f-4d26-4e5c-9eeb-a1b405a74090 · outbound

This paper cites Attentional selection in a cocktail party environment can be decoded from single- trial eeg.Cerebral cortex, 25(7):1697–1706,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Attentional selection in a cocktail party environment can be decoded from single- trial eeg.Cerebral cortex, 25(7):1697–1706,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.240052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.586298Z digest=sha256:f6feccefd90e4ef706bbd1dfd7d03f4c4250de7e221664e47e414f43e12f421b

Observation 6046fe9f-d1f6-4977-8a76-b8d3c9047317 · outbound

This paper cites Neural decoding of at- tentional selection in multi-speaker environments without access to clean sources.Journal of neural engineering, 14(5):056001,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neural decoding of at- tentional selection in multi-speaker environments without access to clean sources.Journal of neural engineering, 14(5):056001,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.058858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.677988Z digest=sha256:99158199f31d0c53c0218adcd8ddf3444b64c265d02c3108d64a3aca05b1f583

Observation b3e6997e-0ba2-4970-ac42-73566d32e77a · outbound

This paper cites Neu- roheed+: Improving neuro-steered speaker extraction with joint auditory attention detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neu- roheed+: Improving neuro-steered speaker extraction with joint auditory attention detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.870687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.758386Z digest=sha256:0312465c4f05312237b6eb0e70244d8d0af252361f6a2e63673bbd5b47bebd28

Observation 20c99bcf-0126-4174-a528-1a75a08b5db6 · outbound

This paper cites Tf-nsse: A time–frequency domain neuro-steered speaker extractor.Applied Acous- tics, 211:109519,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Tf-nsse: A time–frequency domain neuro-steered speaker extractor.Applied Acous- tics, 211:109519,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.706750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.857130Z digest=sha256:3f8e8d6ebb1fcfefaec4d1abf04c2e2df9d5b3614f3c9d15810bc66fd8aed33e

Observation 6acfd61d-a7aa-40dc-a6c3-b881a3eb6a0e · outbound

This paper cites Phase space graph convolutional network for chaotic time series learning.IEEE Transactions on Industrial Informat- ics,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Phase space graph convolutional network for chaotic time series learning.IEEE Transactions on Industrial Informat- ics,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.492949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.950800Z digest=sha256:895d967d773ff2fceffd8efc546339ab203dea0aca21c84db13759c1564d1c1e

Observation 4beca81b-dc3b-475c-aa27-943dd8141bc5 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.286102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.035539Z digest=sha256:a79c7cb94af85f3b076c3c4c3c133305261534021932c9e12490232c199d039c

Observation 8e2a4f39-bf0f-4ec0-ae23-f64f8f361103 · outbound

This paper cites An algorithm for intelligibil- ity prediction of time–frequency weighted noisy speech.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction An algorithm for intelligibil- ity prediction of time–frequency weighted noisy speech

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.136099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.277497Z digest=sha256:2354585f2a1d3a26b0925a433217d76219f116d2dd0c36ddbe9dd18227317186

Observation 3cb4f928-39ea-4f02-acbd-1f4a5fbd31e0 · outbound

This paper cites A study of multichannel spatiotemporal features and knowledge distillation on robust target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction A study of multichannel spatiotemporal features and knowledge distillation on robust target speaker extraction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.950548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.447655Z digest=sha256:0da0755ffd784790676172dd54d8a0663f7e1dcd413e4e64e6231043ef01128f

Observation ff249c62-b548-44b8-aa3d-a2595cb51d0a · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network.IEEE/ACM transactions on audio, speech, and language processing, 28:1370–1384,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Spex: Multi-scale time domain speaker extraction network.IEEE/ACM transactions on audio, speech, and language processing, 28:1370–1384,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.820826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.503332Z digest=sha256:d1621dd54824b4d0678de0e616a7e5544e5465fb53ec97f9e52f9bdcd37f03cc

Observation 3ad76a10-e7c1-4e4d-8d48-72455b9e78c3 · outbound

This paper cites DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:17.955113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.613996Z digest=sha256:657624f49352d334f87a0f90df8225588881c1edc96cc9d3f79e3345125ddd23

Observation bc8e5ec2-d71f-41e9-8950-eb2003d6de73 · outbound

This paper cites Basen: Time-domain brain-assisted speech enhancement network with convolutional cross at- tention in multi-talker conditions.Interspeech 2023,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Basen: Time-domain brain-assisted speech enhancement network with convolutional cross at- tention in multi-talker conditions.Interspeech 2023,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.705495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.692399Z digest=sha256:c9e21efcc9a627cc48f5dd86be48fe9e55a44743c80ee0ee92bfc6962dc3686f

Observation 4c96733a-00b6-4f10-8b93-6354c97f32d8 · outbound

This paper cites Based on audio-video evoked auditory attention detection electroencephalogram dataset.Journal of Tsinghua University (Science and Tech- nology), 64(11):1919–1926,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Based on audio-video evoked auditory attention detection electroencephalogram dataset.Journal of Tsinghua University (Science and Tech- nology), 64(11):1919–1926,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.568754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.749495Z digest=sha256:2ddcb61c1e40081f92e4887a177fe92ac45379b8dd440d953de2af37a3d4933b

Observation 54d544ed-48a9-4f5b-b627-dfc15aff88a2 · outbound

This paper cites Neural target speech extraction: An overview.IEEE Signal Processing Magazine, 40(3):8–29, 2023.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neural target speech extraction: An overview.IEEE Signal Processing Magazine, 40(3):8–29, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.404454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.836014Z digest=sha256:55957618801fef60bb388d3a3fdaeb00d827fb61d8ebc0d3177a7ed48156f9d1

Observation cb6d813a-1da5-4dd1-82c2-343b6f83e2c9 · outbound

This paper cites GroupMamba: Efficient Group-Based Visual State Space Model.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction GroupMamba: Efficient Group-Based Visual State Space Model

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.151270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.151270Z digest=sha256:32cc384214deead2d21c445474d4ca7a40442f3e0b5c75a13536dfe6ff6f6103

Observation c4fc0c39-1540-4c36-83cd-f0c5369100b8 · outbound

This paper cites Centroid estimation with transformer-based speaker embedder for robust target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Centroid estimation with transformer-based speaker embedder for robust target speaker extraction

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.610268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.440078Z digest=sha256:557ef011281e0875fef8761a8a6d11f116cfc46b7825be64fd6086ded9eaeef4

Observation cc40db08-4309-4add-a24d-7547d36f37ff · outbound

This paper cites Speech intelligibility predicted from neural entrainment of the speech envelope.Journal of the Association for Re- search in Otolaryngology, 19:181–191,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speech intelligibility predicted from neural entrainment of the speech envelope.Journal of the Association for Re- search in Otolaryngology, 19:181–191,

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.047429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:17.369243Z digest=sha256:88347ef551aa9553979104de0306c8d7472e24de0d1b85670e39453dc18d23d8

Observation 187771db-3a3e-4b65-a4cd-ec87fac5dad5 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256– 1266,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256– 1266,

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.818684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:16.284920Z digest=sha256:33bbc8a4e92aa56d500abd43276dc9106d3790f1640db020dde7743d8cc3926d

Observation 493b7a6c-6905-47f2-88dd-dccd89182d9c · outbound

This paper cites Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception.NeuroImage, 223:117282,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception.NeuroImage, 223:117282,

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.313507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:14.335409Z digest=sha256:c394fbadd41cc66d408a4e9f49427f624d2fc9dd4e04640e485b068292309e0d

Observation d0c08969-5790-4ea4-b3f3-29de1790de9c · outbound

This paper cites Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.107005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.107005Z digest=sha256:be6454dbb90d630f7ff2d04544a4cfe5962df3d86e93373d3b92f74c08b02967

Observation 0cd41983-8063-42d2-adc1-94a39ac04e1c · outbound

This paper cites NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.245325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:14.465323Z digest=sha256:c6e8885f6387059097e8deec58a28942251c1cec49451e8a30329c062014dd01

Observation f3f7cfbf-b1ee-46ac-8306-2586aa49a79c · outbound

This paper cites End-to-end brain-driven speech enhance- ment in multi-talker conditions.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1718– 1733,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction End-to-end brain-driven speech enhance- ment in multi-talker conditions.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1718– 1733,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.213036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.638340Z digest=sha256:1351e9a388b5b155212ee112a7e6ab6ec7aad0a670efd2ec27a664b6bbdb5153

Observation 65930c91-b637-452e-9df8-32bf922a2c89 · outbound

This paper cites Speaker-independent auditory attention decoding with- out access to clean speech sources.Science advances, 5(5):eaav6134,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speaker-independent auditory attention decoding with- out access to clean speech sources.Science advances, 5(5):eaav6134,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.495225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.015545Z digest=sha256:160ebae862a6271380abef98248aa97e9586bfb786ac49438530f40784c0f11d

Observation f78d3635-634d-4020-9633-6f9d2e4ad98c · outbound

This paper cites X-tf-gridnet: A time–frequency domain target speaker extraction network with adaptive speaker embed- ding fusion.Information Fusion, 112:102550,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction X-tf-gridnet: A time–frequency domain target speaker extraction network with adaptive speaker embed- ding fusion.Information Fusion, 112:102550,

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.157869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:15.228471Z digest=sha256:26f43a99cb58129ec3150616c0d36ac9400a0fac4213e015fdf163546907e263

Observation e8195b1e-72b8-4d7c-aa33-933a0806b837 · outbound

This paper cites Msfnet: Multi-scale fusion net- work for brain-controlled speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Msfnet: Multi-scale fusion net- work for brain-controlled speaker extraction

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.082962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:08:14.580915Z digest=sha256:0277874b346a2ca65d0ef68e5dd7df3f83d3525d2545d897cd7926f27dc5e557

Observation b3fb19f1-d6c2-4a3c-ad58-871fcf6c22aa · outbound

This paper cites SpEx+: A Complete Time Domain Speaker Extraction Network.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction SpEx+: A Complete Time Domain Speaker Extraction Network

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:14.833000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:14.833000Z digest=sha256:edf1269192d52abea7f4b96bc18d6578a88f5be319130b1b0cf3645ad3a4c453

Pith citing papers

Observation 8de67a4d-dd98-4d1f-9a7e-c8cfb58efa2a · inbound

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction cites this paper.

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:41:18.762933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:41:16.865462Z digest=sha256:7e89a6d58470d101c86e576224e83832e25cee2aeb68e83307fc79e04e2c57a1

Observation 36a42793-abee-433f-8157-6c307e1077df · inbound

Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss cites this paper.

Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T07:24:44.110448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:24:44.110448Z digest=sha256:6f536e7d6715cc0f7ef008a63ffa5ff9379f47588f1069ad00999ee1e67557bb