Pith. sign in

Paper Citation Record · LEDGER

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2501.01518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01518 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:30:41.135038Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy42
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ab3960e-a3fe-479b-b16c-aebb3609c7a2 · outbound

This paper cites The conversation: Deep audio-visual speech enhance- ment.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation The conversation: Deep audio-visual speech enhance- ment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.824149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.906332Z digest=sha256:ddd6f3830b9fbd92dcde582d17fdef739f29e711c1992a387b30905cba502288

Observation b5d5e497-055a-4040-86ad-48fa0db08193 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation LRS3-TED: a large-scale dataset for visual speech recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.911436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.911436Z digest=sha256:b05b5acbe304afb207bff15baec2c871f3f4a866ed3982a7cabef939e442736a

Observation 38034054-41b9-483e-9656-bf421910de83 · outbound

This paper cites My lips are concealed: Audio-visual speech enhance- ment through obstructions.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation My lips are concealed: Audio-visual speech enhance- ment through obstructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.813131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.915374Z digest=sha256:eb386f18fe1fb25be1f9d21bf60666df6a971cbc2cb59e80e89573bf9441d62b

Observation d99466b3-8a37-406c-aa62-f2b7d2139348 · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Self-supervised learning of audio-visual objects from video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.802130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.919387Z digest=sha256:26959fc540f87e150c5bc0a85db9a8dfa6cb76f5e756030d3737c82cc7baac7e

Observation db76f7ff-7fb8-4c11-b568-f39062179782 · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation LipNet: End-to-End Sentence-level Lipreading

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.923188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.923188Z digest=sha256:ed6befa1474a141a0ae6136f0098445a75b35bf9ae2d92a13221e8228d55cbe2

Observation 9b8835f8-a4dd-48dd-b9dc-840c56951f08 · outbound

This paper cites Phonemizer: Text to phones transcription for multiple languages in python.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Phonemizer: Text to phones transcription for multiple languages in python

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.791554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.927322Z digest=sha256:42c2b83cae60259c084b33ad806ae9f9ac7b2ca68dd31babe968558a56cca7ac

Observation fd460ee0-f4d2-429f-936d-921a0843a975 · outbound

This paper cites Audio-visual synchronisation in the wild.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audio-visual synchronisation in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.780893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.931502Z digest=sha256:9875121d60adcfee6d58c9fee16308bfa5b68a23e018e4ea74bded277e9cb6b9

Observation afafdaec-dce2-499d-9b34-ac6934cf9b75 · outbound

This paper cites Deep attrac- tor network for single-microphone speaker separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Deep attrac- tor network for single-microphone speaker separation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.770243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.934940Z digest=sha256:d9d20be1537dae58e3019c456e0e93f2d2bd546715b9abd0b10327ea37553d18

Observation 027513d9-d655-4d13-96b1-1b3ddbb520d1 · outbound

This paper cites Lip reading sentences in the wild.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Lip reading sentences in the wild

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.759334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.938885Z digest=sha256:4786e88f59f8bfdc1a5b3c576434a01bfd5878e99f7a7cc33f9925604d8cfc71

Observation dcd99952-8200-4af1-b432-b440c346d91a · outbound

This paper cites Lip reading in the wild.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Lip reading in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.749567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.942276Z digest=sha256:8e092b8fcde2f06d3de80a05852277426ae3137b4e9e0126a26eec4bd12279e8

Observation 1f66ee20-3968-4ac0-9620-6a4d2d0c9e1f · outbound

This paper cites FaceFilter: Audio-visual speech separation using still images.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation FaceFilter: Audio-visual speech separation using still images

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.946090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.946090Z digest=sha256:0625b8fe27f2cb1cc274f8241eaf0c094c1cb7fc6cb69fbfc49ba27b20d7ca6e

Observation 57efc268-dfbc-4605-97cc-bc7a7ed3d7bf · outbound

This paper cites Real time speech enhancement in the waveform domain.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Real time speech enhancement in the waveform domain

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.738937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.949963Z digest=sha256:297ab662007b88b473d0a03661f219554933035542a73d4444da02306e8f8bd2

Observation 7181ccb4-2d5e-4513-bc3d-2092414e194a · outbound

This paper cites Music Source Separation in the Waveform Domain.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Music Source Separation in the Waveform Domain

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.953728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.953728Z digest=sha256:7b8a45fbb3ce7a9662962ab98b4ef4d80eb5c5c3e836fce199f1a9d213f96e3d

Observation b2f14a6a-a786-4146-88cb-40414f437063 · outbound

This paper cites Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.957623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.957623Z digest=sha256:34339d153638ffdae0e95c3956deddb50c9cfbfb43369e8372a33e18f00ca59a

Observation 4b307e38-25a9-40ea-bee5-f689fd46e2ba · outbound

This paper cites Learning joint statistical models for audio- visual fusion and segregation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Learning joint statistical models for audio- visual fusion and segregation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.726885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.961532Z digest=sha256:234ed800601afeee17f4079f4591955ab6a2d0d5403f8911aa8cfe99a3697dae

Observation 9a87c14a-fab0-42d1-b323-d0f98a8f47da · outbound

This paper cites Seeing through noise: Visually driven speaker separa- tion and enhancement.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Seeing through noise: Visually driven speaker separa- tion and enhancement

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.715075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.965596Z digest=sha256:32dfffa3e715c2082fcd5fbe6c61e125939f9f30d9b042043c1ad319a090c590

Observation 14e843c0-fdcf-4304-bb15-488a77c03f41 · outbound

This paper cites Visual Speech Enhancement.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Visual Speech Enhancement

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.969620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.969620Z digest=sha256:ef1755297db14fdbcddedb919d83c2c1d92e862fae65b5a853775ec40ae173ee

Observation 47642839-a43a-4d17-b1c7-4f6bd9bfce5d · outbound

This paper cites Tenen- baum, and Antonio Torralba.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Tenen- baum, and Antonio Torralba

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.703401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.973935Z digest=sha256:861e04c2daa0f79f3d623b32b0c3a532b8ac0243a1b17a3a94245fe568ff3d2e

Observation af3765e5-7ee4-4e70-be0b-70631630a711 · outbound

This paper cites Learning to Separate Object Sounds by Watching Unlabeled Video.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Learning to Separate Object Sounds by Watching Unlabeled Video

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:30:41.243456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.977615Z digest=sha256:24f54a7f6ea16ebc84f981230036edb76d8d210a7965648b0cccc7b6dd3798da

Observation 043b39e3-2d44-46ed-aec7-e68bb17ab545 · outbound

This paper cites Co-Separating Sounds of Visual Objects.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Co-Separating Sounds of Visual Objects

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:30:41.228155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.981589Z digest=sha256:f514034392d5e98b75f65ba568e45a1841c145b10b1e8d198411d61332e0560a

Observation ea8cc40b-b7da-4e9e-a01d-f9f7e666a425 · outbound

This paper cites VisualV oice: Audio- Visual Speech Separation with Cross-Modal Consistency.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation VisualV oice: Audio- Visual Speech Separation with Cross-Modal Consistency

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.692101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.985549Z digest=sha256:852e92faf873d7c1b3fa706ef7ec3d0382437d4bfb4b97fa72ec0b14b7926819

Observation 95c0dff2-817e-4b81-a01c-9447e9188322 · outbound

This paper cites Multi-modal multi-channel tar- get speech separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Multi-modal multi-channel tar- get speech separation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.681770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.989309Z digest=sha256:2461bb6d156a3eeaa067d8022cb61f5d78b3732257a07b0d390714fb673f1808

Observation 895d562e-4cf0-488d-ac62-931e5016a1db · outbound

This paper cites Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.670976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:40.993004Z digest=sha256:8ffef1fe015e844e67a7d32d1c2285b3955342afe1b5a8b76f3061169ff23f78

Observation a6536841-966a-446e-9df7-b7a1a40b160d · outbound

This paper cites Perceiver: General Perception with Iterative Attention.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Perceiver: General Perception with Iterative Attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.996705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.996705Z digest=sha256:8587ca30e82942b36bb4032d67ff25ecf88381d2ae896e8f60235ed77126842a

Observation 29573d01-f545-40e5-97fa-0338a57339cb · outbound

This paper cites Looking into your speech: Learning cross-modal affinity for audio-visual speech sepa- ration.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Looking into your speech: Learning cross-modal affinity for audio-visual speech sepa- ration

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.659982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.000900Z digest=sha256:47ea47cda119c9febacc491813b26e9618e43e05fa7ae5f1e842fa6683f3a97c

Observation b5081915-2b49-492b-be21-8e7cefaff36f · outbound

This paper cites Parameter efficient multimodal trans- formers for video representation learning.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Parameter efficient multimodal trans- formers for video representation learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.650080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.005406Z digest=sha256:3c8daa841bb8919857c25f06cd47ccb741b2f4d89d96a5c2f740ba91638cc6e0

Observation 16cd0da0-dcf0-4149-be3f-f141833517bd · outbound

This paper cites Audiovisual trans- former with instance attention for audio-visual event local- ization.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audiovisual trans- former with instance attention for audio-visual event local- ization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.638710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.009871Z digest=sha256:5dafa7103b1ce15594b22226ba6d0d59c2c929306e1a8733a8b0f825a4931615

Observation daeb5589-62dd-47b6-97e2-52ef461facb0 · outbound

This paper cites Speaker- independent speech separation with deep attractor network.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Speaker- independent speech separation with deep attractor network

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.627331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.013675Z digest=sha256:de2bb06a8db182fc01ad225ccd02d78268beeaa2fe6e5e70a562fd260d33db9c

Observation fa8a0789-3ef8-4ad0-b9fa-870f2b0d0139 · outbound

This paper cites Attention in dichotic listening: Affective cues and the influence of instructions.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Attention in dichotic listening: Affective cues and the influence of instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.614029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.017791Z digest=sha256:ae2a50cad573e887116392821c270745a52b0d9c3b34aad0743a0949aca78b46

Observation 0027f422-4a53-40ae-83c9-28e7c7e097e5 · outbound

This paper cites Attention bottlenecks for multimodal fusion.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Attention bottlenecks for multimodal fusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.602676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.021525Z digest=sha256:84674e309348321b298f7b5a4a583ca13ae3a130471ea1bae747a4c3ad1b7aa7

Observation 96797977-ada7-4f82-9afd-12465ab51e46 · outbound

This paper cites an unresolved cited work.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:30:41.591631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.026076Z digest=sha256:6ad845d63ab1382372347e3a8b65480db02aec6a11ed4ecda81a8a117785f61a

Observation 5dbec9bf-5987-4585-93b1-1f46c80d6a0b · outbound

This paper cites an unresolved cited work.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:30:41.580932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.031979Z digest=sha256:911333fcd26872bddc0252643a16ce6db73a9a7fe1cc9c3b36746a4c02b34a23

Observation 1b99f204-5e96-45ae-be0d-f747a5de7cff · outbound

This paper cites Audio-visual object localization and separation using low- rank and sparsity.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audio-visual object localization and separation using low- rank and sparsity

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.569605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.035764Z digest=sha256:efa008b79c2a03cc7414e92a6e732352b9ac3400b11cefc79d5ca192bd8b5247

Observation 2ab1403e-7f12-4380-87af-c1df667eb2ee · outbound

This paper cites mir eval: A transparent implementation of common mir metrics.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation mir eval: A transparent implementation of common mir metrics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.558567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.039549Z digest=sha256:e989f8e7ced2248de281341eec8a2f8e2bc4a4376e097724ca74f73bf3f26055

Observation f89505bb-77e7-4f4f-913d-95f3628267a9 · outbound

This paper cites Interspeech 2021 deep noise suppression challenge.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Interspeech 2021 deep noise suppression challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.547587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.043823Z digest=sha256:dfaa0c5070f2713078c4185f2236b6957568a7a327e26f17d9b7be884ba8a7f9

Observation 1605df20-c4fb-4ef3-9265-94cc66d4dc32 · outbound

This paper cites Visual keyword spotting with attention.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Visual keyword spotting with attention

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.535480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.048654Z digest=sha256:4a08b8233992ca13ac2221ac9fb611b84aa19b847539ced0149f86b28de5b95e

Observation 98ba1576-386b-4233-8484-dee668835509 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of tele- phone networks and codecs.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of tele- phone networks and codecs

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.523260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.052999Z digest=sha256:e44dfac1d81dbf7e9f9caa63825d91f2c2c3219df3b64b81ea9ddc595632a324

Observation 7bf38965-c9ea-4739-9ec7-3dfa4869b105 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation U- net: Convolutional networks for biomedical image segmen- tation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:41.056750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:41.056750Z digest=sha256:26f395dfa773b6e5baba7f73621c720ba8deaff4b9fd5494681945f9c06f04b1

Observation b46a5225-02d9-431c-8c60-1da3672a0ab8 · outbound

This paper cites Self-supervised audio-visual co-segmentation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Self-supervised audio-visual co-segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.504725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.060530Z digest=sha256:56cfac294f1cfae0f5d562aa47bc59a41a4bbddf30ae90d94ce7801ee4abd3e3

Observation 8439aeb5-1892-4af5-93e5-8858be541a88 · outbound

This paper cites Audio-visual speech enhancement using conditional variational auto-encoders.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audio-visual speech enhancement using conditional variational auto-encoders

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.493571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.064735Z digest=sha256:da9959cd808aff8093fd29eab42d3a95e5e382c29d8e2b52ed393783ee5e0e62

Observation 8958c87f-7317-48ca-8abc-2a206c2c2e68 · outbound

This paper cites Seeing to hear better: evidence for early audio- visual interactions in speech identification.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Seeing to hear better: evidence for early audio- visual interactions in speech identification

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.482804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.068465Z digest=sha256:9de669bd9f395032ec87ba088c3990c85f4a613370f641a9496461ce4a032f0f

Observation 925342c2-ae96-4f7e-b01b-7464fd4eef87 · outbound

This paper cites Combining residual networks with lstms for lipreading.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Combining residual networks with lstms for lipreading

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.471859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.072889Z digest=sha256:dcf62ad41d1979a2e745b585c39d0fd5789aca5c9bf7cc6e69e143bcf82b7a01

Observation f5d63be5-f809-4ebc-8f76-4623a4669aa5 · outbound

This paper cites An algorithm for intelligibility prediction of time- frequency weighted noisy speech.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation An algorithm for intelligibility prediction of time- frequency weighted noisy speech

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.459504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.076646Z digest=sha256:64ea2ae9ae559267d164786120e55762dee4781d8c5107ffc3b3520ad40688de

Observation bf6c4257-0306-4ed4-b4a4-53236c4e19cb · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.448453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.080374Z digest=sha256:f32eaf762b76614da6833761ca8231f71d2ce286fe27df3278085c624c83766f

Observation 01094ad2-422e-4d84-b36a-b305d6649124 · outbound

This paper cites Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:41.084566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:41.084566Z digest=sha256:295da4fca92405223b3c3459ebaa9d054e041a3bdd67e973dae54612fd256330

Observation 70ce8276-f361-4f3d-a88e-9a3837b1bd3c · outbound

This paper cites Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:30:41.188958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.088534Z digest=sha256:a9e705ad98a665b6d4885fffa7a8dbddf4974aaf26a0f855ccbfb2c961ece4f0

Observation 6c9a8694-407f-4356-aad8-6d1fba8e3278 · outbound

This paper cites BSS EV AL toolbox user guide.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation BSS EV AL toolbox user guide

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.437238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.092693Z digest=sha256:2df205952e2bd77539e22a0edb20a72a476c74007eb38e70ed4cd164cf9129c4

Observation c4a5be8b-1a04-4b42-8acf-601a24046850 · outbound

This paper cites Supervised speech sep- aration based on deep learning: An overview.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Supervised speech sep- aration based on deep learning: An overview

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.413067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.100737Z digest=sha256:1015d6ffab9381b1b5e0fc1a2fa7f8d370cdf95175df804f3e5cc9d6f302611b

Observation b40b4959-5bfd-4f45-bca7-24f196a58ba5 · outbound

This paper cites V oicefilter: Tar- geted voice separation by speaker-conditioned spectrogram masking.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation V oicefilter: Tar- geted voice separation by speaker-conditioned spectrogram masking

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.401200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.104155Z digest=sha256:84e97c545ffc8b7f20e246e424d2cf88d74ace6a5f43baa760275c75f0cd8979

Observation 48bcb8d7-55c5-4300-a1c9-f683f67eb681 · outbound

This paper cites Combining spectral and spatial features for deep learning based blind speaker separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Combining spectral and spatial features for deep learning based blind speaker separation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.389835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.108114Z digest=sha256:3b128dcdebbd7f3999d3453110edcf9be2c301678f1861bcacd91a74a1c94ad6

Observation aec89adc-0dbb-4aa4-80d4-390ae8562903 · outbound

This paper cites Time domain audio visual speech separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Time domain audio visual speech separation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.377627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.111482Z digest=sha256:8761e73e3d01e2ab7cebee2a2864229af29f5c742428d9f3e0406a8ae3dac0e7

Observation 46297670-87a4-4164-844e-89649023ac2c · outbound

This paper cites Multilevel language and vision integration for text-to-clip retrieval.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Multilevel language and vision integration for text-to-clip retrieval

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.365307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.115377Z digest=sha256:6b927dcdc35757a0e90f5eb6543a3877201ed12e39e5f0f662b695bf061adef9

Observation 7c021b59-cba2-4afd-b9c2-b8154d58eb59 · outbound

This paper cites Recursive visual sound separation using minus-plus net.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Recursive visual sound separation using minus-plus net

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.354847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.118996Z digest=sha256:fa879534c0c662601c8adf438c10da1742acfc3a0e7a843e3a42f724fe11ef3c

Observation bd574924-a41b-45cd-87e1-d1041bfb2cfe · outbound

This paper cites Permutation invariant training of deep models for speaker-independent multi-talker speech separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Permutation invariant training of deep models for speaker-independent multi-talker speech separation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.344301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.123101Z digest=sha256:bce74079302ec86f4b5c1d29e9a0d4c079a16d6c89b84127f3743e2fd16f8e0e

Observation 3a19ff94-23a7-4a91-b9ef-9504a7120d69 · outbound

This paper cites To find where you talk: Temporal sentence localization in video with attention based location regression.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation To find where you talk: Temporal sentence localization in video with attention based location regression

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.332912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.127305Z digest=sha256:9af521a2dd726c1e0a1956c1be3554d8a42a814c4fb168106aa732bb8f878e9c

Observation 874a4fa0-1153-49f4-9d49-67239dd48ae4 · outbound

This paper cites The sound of motions.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation The sound of motions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.321307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.131259Z digest=sha256:859d24e5ab71695201840b61670d9bc9abaf1d867541400ba428808c1bf58781

Observation de3c0aa3-f74f-45a9-9044-40555c7f02a1 · outbound

This paper cites The Sound of Pixels.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation The Sound of Pixels

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:41.135038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:41.135038Z digest=sha256:0c42cbb72494a4adac9f481116ddaefc613d7d391d787b4634d7837c9ccc7141

Observation 6aec83be-3d62-49aa-9182-461b5a26161e · outbound

This paper cites an unresolved cited work.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unresolved cited work

Reference 1706

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:30:41.425188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:30:41.097036Z digest=sha256:85420c85fbe81188002a631cea540057bff707aaf070bef62df71c885dd97bc1

Pith citing papers

No inbound Pith citation observations are available.