Pith. sign in

Paper Citation Record · LEDGER

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

As of 14 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2608.09288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09288 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:17.649998Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:17.344737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T20:11:18.008189Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc445400-b253-469c-9d90-86a84a66ae2c · outbound

This paper cites an unresolved cited work.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:11:18.851075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.333888Z digest=sha256:446accee73c040cc646f2a922fb2c24f3ae19eb85d9fe6afb374c905887671fa

Observation 25cdecb6-faa7-4f86-8dc5-9fec91c89c82 · outbound

This paper cites DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-11T20:11:18.015625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.344737Z digest=sha256:c64d9e19f82cf2c1191d81fdaaa1fbe59cfd021e09f21eca484bba52eae41f7c

Observation a91461eb-53fb-4c65-a36f-5604f41ea4e0 · outbound

This paper cites As illustrated in Figure 1, the framework comprises three components.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation As illustrated in Figure 1, the framework comprises three components

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.829440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.354145Z digest=sha256:0d47a67112ce6dd3616a558c81b7b72f268f4f302bb56e4549898c9a33fb4264

Observation 66e2904c-72a7-44cf-942e-09160a56e2a8 · outbound

This paper cites Each GPU processes a batch of 2 three-second segments.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Each GPU processes a batch of 2 three-second segments

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T20:11:17.979671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.363813Z digest=sha256:18779a607db170f62b26a1e371141884698c706ced092870d1dc814cc4350fb8

Observation 75eea7ea-8709-4388-8bb1-1e3ec166f98e · outbound

This paper cites an unresolved cited work.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:11:18.809126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.375982Z digest=sha256:980b1dd638fcb99f57cde589a9cb5b874251883a1254fdeddcf81ab70dbbdf56

Observation 83323de6-c4ea-426e-952e-d990f5845e29 · outbound

This paper cites Conv-TasNet: Surpassing ideal time- frequency magnitude masking for speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Conv-TasNet: Surpassing ideal time- frequency magnitude masking for speech separation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.787881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.384365Z digest=sha256:3eab6e6e67f64099a517858b542ca5f1eb70e713515cd69fb679ce97287f1159

Observation 115e74d8-e6ce-45f5-a227-59b676efcbc3 · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.392252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.392252Z digest=sha256:aa299d7eeae7c74c89a0c963547b0b4ae5cc8093c844399935b4005638f2e7ce

Observation e8a1fcb8-2e8a-470c-b1f9-6c4cf83f3fd5 · outbound

This paper cites Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.401926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.401926Z digest=sha256:568ba0cd06f5b56e4de409b252754541697c093cbd4dce3ea88d22a149701b93

Observation 57262555-fa9a-4c87-9042-8815d639a69f · outbound

This paper cites Ps4: Proxy-supervised joint training for real target speaker extraction,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Ps4: Proxy-supervised joint training for real target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.733292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.408624Z digest=sha256:cff5814b815d5797cee177ae14b64b4088e0c796b261b0daac4de9b38020512f

Observation 6bb8351a-2688-4b3f-a87f-a37375ae6694 · outbound

This paper cites Attention is all you need in speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Attention is all you need in speech separation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.468283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.468283Z digest=sha256:12435355a9ce28d1da341a43fdb42fc57a854f03718c51ae1829604489107532

Observation 5e71a295-8a89-476d-bb03-6856ba402afb · outbound

This paper cites Looking to listen at the cock- tail party: A speaker-independent audio-visual model for speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Looking to listen at the cock- tail party: A speaker-independent audio-visual model for speech separation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.712514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.426200Z digest=sha256:7dfba1a367ddbb80f16e3d3ecf443b5f611c6e456f4afce5e9225e518f9407c9

Observation 96b59652-e80d-4c7c-8076-0f3dc98bea56 · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation The conversation: Deep audio-visual speech enhancement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.693661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.436060Z digest=sha256:aafef555ec95114921e18b60723c00f5d1a90abf027ca0e9e34f18c1fb724561

Observation 992753b2-b15a-4bfb-b1e8-9c9bbfe2381a · outbound

This paper cites The fifth CHiME speech separation and recogni- tion challenge: Dataset, task and baselines,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation The fifth CHiME speech separation and recogni- tion challenge: Dataset, task and baselines,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.673732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.447103Z digest=sha256:daf3934652d1f33877d2571600c52870c09b8f566225ecaf5e255e69a18d87e2

Observation 3f8a515a-c3a2-4c16-84e0-966d748376b7 · outbound

This paper cites Time domain audio visual speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Time domain audio visual speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.653153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.455231Z digest=sha256:733e69251fed7d01da76e8b0a849ac1af75db8a7731c30cf1463e06ae200c66f

Observation d71aea48-a37b-45d7-8ae9-f0a815ffa552 · outbound

This paper cites Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.631428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.461553Z digest=sha256:3531f8261362f61e58979d91bae78d7a1ec3e01b18b4be3979e6bd70b21e4f58

Observation 253453a5-10ff-4d60-a783-d704dc62d06e · outbound

This paper cites Tiger: Time-frequency interleaved gain extraction and reconstruction for efficient speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Tiger: Time-frequency interleaved gain extraction and reconstruction for efficient speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.443078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.522082Z digest=sha256:fc9a5966fcd993d106ef31ebebdc7036108b7ec5cda4b89ee2e0f84a53606a76

Observation d2608980-c4e3-496e-b4f6-25dfff961332 · outbound

This paper cites The AliMeeting corpus: A multi-modal meeting corpus,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation The AliMeeting corpus: A multi-modal meeting corpus,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.595219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.473940Z digest=sha256:eed3fb3e3833e0f7242ddfec5ec75024b11b69f36c1aa29230a030c53b702106

Observation fd809e29-c043-4c39-9c87-f7b69a1723b6 · outbound

This paper cites MISP: A multi-modal interactive speech process- ing system,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation MISP: A multi-modal interactive speech process- ing system,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.572367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.483059Z digest=sha256:daf8d8d49ab04601f6fece6cefebe957b1fc412f3da19f0ff8b253c21ad65b0f

Observation 3550d6bb-89c4-4abc-b4b9-12293e97788f · outbound

This paper cites AISHELL-4: An open source dataset for speech separation, diarization and recognition,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation AISHELL-4: An open source dataset for speech separation, diarization and recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.549663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.494065Z digest=sha256:55a821890f34a2c299e7b0697fb5aa12e2d16d0b8f55d432c2eff8340ddacb28

Observation 7d840653-aa0e-466d-96f8-e7feaa5005db · outbound

This paper cites Pyroomacoustics: A Python package for audio room simulation and array processing algorithms,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Pyroomacoustics: A Python package for audio room simulation and array processing algorithms,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.517409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.501409Z digest=sha256:b710c6b1adb56ba4a9438aa12a7126861d5fa667c237671bb11158fc79c32363

Observation 493c7620-d86e-44ec-b7c9-abc3c55dfb05 · outbound

This paper cites A tutorial on hidden markov models and selected applications in speech recognition,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation A tutorial on hidden markov models and selected applications in speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.485875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.508585Z digest=sha256:f161bd51fe9348d9e7eba766aba61f9d19c826827e81e055acc33d9fa2cccb53

Observation 17111b39-8046-4a0f-80e2-243022073d20 · outbound

This paper cites WeSpeaker: A research and production oriented speaker embedding learning toolkit,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation WeSpeaker: A research and production oriented speaker embedding learning toolkit,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.324783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.570686Z digest=sha256:e3246a58ca472fc7059d5047cf8b177b913980a2468449a67323405d73b1ee23

Observation 0091574d-79e4-4869-bf45-747464705955 · outbound

This paper cites FunASR: A fundamental end-to-end speech recog- nition toolkit,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation FunASR: A fundamental end-to-end speech recog- nition toolkit,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.422808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.530326Z digest=sha256:2caee2399b25c551a0ba66c5901d25204d0f6766e2936373d144e0242808cbcd

Observation 10891722-f7dd-4c05-974e-b8140a655b55 · outbound

This paper cites Room impulse response generator,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Room impulse response generator,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.394748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.537504Z digest=sha256:08450df94e72dd16883ad40952e0ea95d71adea2871e729c6e8953bfce9217b5

Observation 4c61cefb-b562-47f4-adbc-271e34d9e0f3 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation MUSAN: A Music, Speech, and Noise Corpus

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.545140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.545140Z digest=sha256:839aa6eb2b78351371b7f1f04e59c9a4913c5a9058051f31b4f7fb15885de5fb

Observation d49fe0c5-3e18-47d9-bfa1-bbf262df550c · outbound

This paper cites Overcoming catastrophic forgetting in neural networks,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Overcoming catastrophic forgetting in neural networks,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.553114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.553114Z digest=sha256:6a6fca531740f95139a99d8d738b64b2180deb45ce440e89309346816a4162a1

Observation cd12754a-dbc8-4030-b37a-449c3f8bbe28 · outbound

This paper cites SDR— half-baked or well done?.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation SDR— half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.350823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.561434Z digest=sha256:88c202b66ed6db6486b6be4b139683a6a40dc02909a3010fbb250c41b5cac654

Observation 80506bb9-9d1d-4262-9005-0f181893d814 · outbound

This paper cites Out of time: automated lip sync in the wild,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Out of time: automated lip sync in the wild,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.182349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.607522Z digest=sha256:98acd9c20f6c7d80394a07a4ee6e230b3d283d92da4e6ba0c174a0c1dd82a5c3

Observation 4621ba01-de89-496c-b84a-3f695fe93e4c · outbound

This paper cites An al- gorithm for intelligibility prediction of time-frequency weighted noisy speech,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation An al- gorithm for intelligibility prediction of time-frequency weighted noisy speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.292973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.576768Z digest=sha256:10e7ec140697d760c2adbc088115ce9b64b846e021f7c6380474fa58095ed553

Observation 37610fde-13b3-4199-b35a-cf321ca2121e · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Per- ceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.259065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.583386Z digest=sha256:fc40bca410718e96a63ea68dba4e67feb77978a3190642572a021aa4e2a3b543

Observation c499a89d-fa14-4693-a02d-f02d1cfabeca · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS challenge 2022,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation UTMOS: UTokyo-SaruLab system for V oiceMOS challenge 2022,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.233862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.589283Z digest=sha256:51fe21ab5ef1087b59082b56678a90caa434875a061af56e91af3f63e21774ad

Observation 38e85c96-5376-4b67-a3d7-376e9a4f4a18 · outbound

This paper cites CN-Celeb: A challenging Chinese speaker recognition dataset,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation CN-Celeb: A challenging Chinese speaker recognition dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.207925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.594834Z digest=sha256:0f7798c616646d7b7c615a18580c3e95d02b30e62e418affab8a2abe8768fd9d

Observation bdaae359-ff6e-41a6-a41f-d9f99bc321c9 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:17.601296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:17.601296Z digest=sha256:e3760626925989bfc5cdf55146e47edb4593067e19ec4072b1564c42beac3645

Observation 98347166-25ef-4caa-be48-4d0cea703896 · outbound

This paper cites DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.056216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.645482Z digest=sha256:74288f79dfa0e5d4d8965cbe88feefaf2524be1d1829982c6e1ee6d7e26faac8

Observation bf281f55-d30e-403c-91c5-025d59aa647a · outbound

This paper cites Perfect match: Improved cross-modal embeddings for audio-visual synchronisa- tion,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Perfect match: Improved cross-modal embeddings for audio-visual synchronisa- tion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.161583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.613825Z digest=sha256:8561ab67e59c9879351a846ad98fa6736c232d65e3780bc3ff9f8b99a51578f8

Observation 09a28137-d12b-41e7-b2de-3002d75bf121 · outbound

This paper cites How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks),.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks),

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.141478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.622982Z digest=sha256:3f4406f744abe150a8676c3cb3623dce2c70553405c1bb0435b8674da8692ac1

Observation 756fc988-95c4-4891-bd19-25e58cd4a8ae · outbound

This paper cites SEGAN: Speech en- hancement generative adversarial network,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation SEGAN: Speech en- hancement generative adversarial network,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.121792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.629274Z digest=sha256:fa0e520d8092bcf093fa892771507aceec5f7ef92a384208050641a584046ee8

Observation 2e55f7fb-f54a-4281-959d-9e2b2fe8e520 · outbound

This paper cites MossFormer2: Combining transformer and RNN-free recurrent network for enhanced time-domain monaural speech separation,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation MossFormer2: Combining transformer and RNN-free recurrent network for enhanced time-domain monaural speech separation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.101171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.636161Z digest=sha256:70eb1d9d58ea0a854bfcbd332f4bd92f041d066f6dc4edf4eaf90c33d9e81205

Observation 37ccb782-15f1-44b6-aa2c-b3cf76884711 · outbound

This paper cites Recommendation ITU- R BS.1770-4: Algorithms to measure audio programme loudness and true-peak audio level,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Recommendation ITU- R BS.1770-4: Algorithms to measure audio programme loudness and true-peak audio level,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.078598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.641112Z digest=sha256:cefa6f4d5da8aa8b70d6bd0fdb7522d25451556e10fc91b344c6e6729099d30c

Observation 58be2944-db1e-452e-a1ce-174982bea642 · outbound

This paper cites Adam: A method for stochastic opti- mization,.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Adam: A method for stochastic opti- mization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:18.037042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.649998Z digest=sha256:6f9e7cb8181b3b31c5d7ef4081b3891182fda9359e96379340d8e8a93040d8b3

Observation 3a2c279c-6dfb-463f-8a7b-c2b2b003c788 · outbound

This paper cites PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T20:11:17.777692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.417789Z digest=sha256:fa8107c270d026d8c5878c95eaed002e0c14d7fea0e6696241f68db038efe63d

Pith citing papers

Observation 25cdecb6-faa7-4f86-8dc5-9fec91c89c82 · inbound

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation cites this paper.

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-11T20:11:18.015625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:11:17.344737Z digest=sha256:c64d9e19f82cf2c1191d81fdaaa1fbe59cfd021e09f21eca484bba52eae41f7c