Pith. sign in

Paper Citation Record · LEDGER

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

As of 22 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2606.23712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23712 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:17:45.299833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T22:49:01.384665Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5deb219c-beb9-49e1-b4f4-5d425e4dabe3 · outbound

This paper cites Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Recent deep learn- ing approaches have significantly improved performance [1–3], with generative modeling frameworks emerging as a powerful direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:f86c96c7189333a6fa06909cdf4c17f36be2824e498e1dd1834248348e224220

Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · outbound

This paper cites Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:49:01.387089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:8eb58de4205ea246d08eeed6db2df5bbdcb4df4a097d9d137c454e858dbe60e2

Observation a9516dbd-c7dc-4cc8-9b84-34a969b732b5 · outbound

This paper cites 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement 1, describing how the audio-visual con- trastive loss is computed and the motivation behind it

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:320a259eec8d72aa3007e7bd1dc0aad8c2c2b8a53eb82ffcfe9044c4fa09cbf3

Observation 29ed98d9-149b-47af-ad22-c702fad92144 · outbound

This paper cites As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7].

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement As baselines, we con- sider A V-DiffUSEEN, with cross-attention fusion [8, 13], the audio-only version, AO-DiffUSEEN [13], and the supervised- generative FlowA VSE model [7]

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:291df44d65f5a3ce195dd6ee539214fca403c7bcb5747541abc8557fc5473c61

Observation 78d6f946-dbcc-498e-a989-35221d479999 · outbound

This paper cites Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Matched condition: TCD-DEMAND Under matched conditions (Table 1), the proposed model im- proves all metrics compared to A V-DiffUSEEN

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:ff8da766042cc214aa23c64374e2ecf9ac97fd2e3e84a57c54917e7feb6725e5

Observation 524e131a-c194-485f-a7c3-9ed5ce3046b2 · outbound

This paper cites During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement During the pretraining of the visual-conditioned speech diffusion model, we augment the denoising score matching ob- jective with a contrastive audio-visual alignment loss

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:c81b09ce0756d4e631c570a1e763e4fdf2accb44f938ceab728cdf7b11540b6c

Observation a8dccfb8-f91f-46f7-82d7-19fd41abd647 · outbound

This paper cites an unresolved cited work.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:8857261a051ec20b79eb7524073510bc6b93ba2625ef49664a0ceb8d9ce7af80

Observation abe0aa72-a4df-4085-a5ff-2157cc113e88 · outbound

This paper cites Generative AI tools were only used to edit and polish some portions of the manuscript.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Generative AI tools were only used to edit and polish some portions of the manuscript

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:41e16ada2245543d0937d4dad1572e2923cfa24bb36f80f4d4cf1732e842bb94

Observation bc98de02-c35e-4519-9557-50233c8b473a · outbound

This paper cites SEGAN: Speech enhancement generative adversarial network,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SEGAN: Speech enhancement generative adversarial network,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:1705e17a1439873f4a3ad9d053f9d4892840ef444ed89812f27a2cc23451c553

Observation 42c0a82b-ce2c-4996-9b44-0c0f8463d353 · outbound

This paper cites Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:08f47646c53599f99287fe5fb4d6354c6b9372f821976693e7b61af83fcb5883

Observation 7836dbb8-7d74-40fb-ae77-af0bb744b29d · outbound

This paper cites DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:a0083f6fef5b6e70b8c140895aabb2f1a26432f74504b90848d821b9b7243cbd

Observation 7e54e6c0-1e8e-4fe8-9274-037b1169f730 · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based gen- erative models,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Speech enhancement and dereverberation with diffusion-based gen- erative models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:c30cf986ced385deaaa731f1f598630561196777bf821ba75835eddcfb1d4ba6

Observation ccded98c-1c52-4c7c-a183-7c800ab26646 · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The conversation: Deep audio-visual speech enhancement,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:ca4e0fe912446587406418d0556fc2f8e5da0330b7fdb475868b11651d8b20a9

Observation 6e168e88-24da-43a7-9283-ded77279913f · outbound

This paper cites Audio-visual speech enhancement using multimodal deep con- volutional neural networks,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual speech enhancement using multimodal deep con- volutional neural networks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:8d3acc19c4efe04263121f34d2b2879c2596c980500655bad1a2a7e04fb09081

Observation 2fa0d510-f391-4a3f-97a8-4915a7be2514 · outbound

This paper cites FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement FlowA VSE: Efficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:a4f70cf1e909d610d065fd58c2e4004a541555c761ca78352a5ce72f10a3a4c8

Observation 0c7e79e7-2730-4bf5-a57d-4f0ba8828a25 · outbound

This paper cites Diffusion-based unsupervised audio-visual speech enhancement,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based unsupervised audio-visual speech enhancement,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:3fd39cc1bc639a428f97e0a0854be622f077e86139708de07653659ea0df9908

Observation 00df5926-de7b-48df-95f8-0b002ed8d94a · outbound

This paper cites An overview of deep-learning-based audio-visual speech enhancement and separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An overview of deep-learning-based audio-visual speech enhancement and separation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b521aef9003a0087a7178fa1f414931c73948c0066248b6f2eeb2a6a96793ac5

Observation 3b129cf7-9d46-4a10-83dc-ebecdac0bbec · outbound

This paper cites Learning transfer- able visual models from natural language supervision,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning transfer- able visual models from natural language supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:035bca7ef7cf030d218519fdf8117ada03ff238b791699e9f50f0c9a8218c688

Observation 3ed6430b-a876-42ef-a7b3-fc3397cc4d8c · outbound

This paper cites CLAP: Learn- ing audio concepts from natural language supervision,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement CLAP: Learn- ing audio concepts from natural language supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:bc2a77de496270d57fd724fbf24cd1b159b4db9aed6cfbfdb71ad9c4be661c0f

Observation 032e5418-8cae-463e-87a7-1922dc297312 · outbound

This paper cites SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SA V-SE: Scene-aware audio-visual speech enhancement with selective state space model,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:5e0e474abf5c5ce531ce799c99a896a64ef0b056b425dbe71855e83c2b651a9c

Observation f9107a07-13aa-4db5-aaa1-dd0767c09fdb · outbound

This paper cites Diffusion-based Frameworks for Unsupervised Speech Enhancement.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Diffusion-based Frameworks for Unsupervised Speech Enhancement

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.372695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:c55e74e9ab0c61f80d6d371eecfb3835f684d59983b100198cd0e37aa47a58cd

Observation 9ec4f8fc-b3ab-4292-9e12-fb2a8417dbf2 · outbound

This paper cites Score-based generative modeling through stochastic differ- ential equations,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Score-based generative modeling through stochastic differ- ential equations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:1f9a015615ce80357fdd71d3c890dc25d875961a878376443f3c02e675598f7d

Observation 4ad42cd0-8bd9-49ef-ac26-3299a6263287 · outbound

This paper cites Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Nonnegative matrix factor- ization with the itakura-saito divergence: With application to music analysis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:ed31af57096db2b3d5bf3a841acddfe50717028ea436b7d8ea6f84ae81c7d919

Observation 4637457c-8421-4014-a872-d69d9cc1cda5 · outbound

This paper cites Tweedie’s formula and selection bias,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Tweedie’s formula and selection bias,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:82d63266013a94b55d98c10de384a155809518ffac1cf1c35b145f7feb700e48

Observation eea11571-501f-4479-acdc-02b3503f7786 · outbound

This paper cites Deep residual learning for im- age recognition,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Deep residual learning for im- age recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:1b307ad68be57383ea4445d47e3ec70582b34d5b78f070003497b27e9e910f1c

Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:01.368464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:2ef6a0a842f47bb4d65323226bacc1f138873f97dfca9a46bcb35e2a964c379d

Observation 7d69518e-b9cd-4010-9d25-aa3d18057d95 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Representation Learning with Contrastive Predictive Coding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.377870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:170b0a42345917f19ba4b098b7ad95e33d27150aa12cbb3df07400debe000d8f

Observation fcbf7481-a563-4ee6-a647-5048a5bc8b76 · outbound

This paper cites TCD-TIMIT: An audio-visual corpus of con- tinuous speech,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement TCD-TIMIT: An audio-visual corpus of con- tinuous speech,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:ade8c5cb6fdff10b686e55ca2c39ea4f94c66cd16f3ac7691dfff84140ca6532

Observation 320a8621-46b7-4b0d-8d90-173f4644e06e · outbound

This paper cites The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement The diverse environments multi- channel acoustic noise database (DEMAND): A database of multi- channel environmental noise recordings,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:bee23b9fe71bc5a70dc82db588f3f329530c79a969620814c77f72e41695c084

Observation 63597d9c-061d-4644-930c-d95d8a30ea95 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement LRS3-TED: a large-scale dataset for visual speech recognition

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:49:01.382024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:17f4926f2c662f26f0c1ce392b0770e62181f3d8f0d1139a3c138a9e59d7de20

Observation 7016b475-36d9-4462-bdd4-24d74b2ec68b · outbound

This paper cites NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement NTCD-TIMIT: A new database and base- line for noise-robust audio-visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:00c6d724cd5173852a1faaee79693d73b57cb11e8862c9ed340cb70ab214b683

Observation 709ebb50-95b1-4e58-b620-921464db7ada · outbound

This paper cites SDR–half- baked or well done?.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement SDR–half- baked or well done?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:5a88a6106f008cf572c901a49c0a4b14dc0cba1be8cba59272fba095b23c042d

Observation bcd495df-920c-485f-85e3-ddb03a86f8fc · outbound

This paper cites Performance measure- ment in blind audio source separation,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Performance measure- ment in blind audio source separation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:b3562e354b69cd58467bdc45ae45069320a8d12fdf492c521483906fbb2ce4b2

Observation 272ef1c0-95a9-4e6f-a1f0-8d00467d0975 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:a329dd4a176b1442d3a3b1102ff0eae51f5e2095a229ce583c1c78ac190b775e

Observation 0983f300-b362-4d03-a785-f905f04a78b2 · outbound

This paper cites An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement An algo- rithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T23:17:45.299833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:c3991d3aa05e5257f7fad84e53f68d4a51b7df1f03b67a23efbaed6794d738c3

Pith citing papers

Observation 89dbd48a-c87c-4bca-acaa-4f190bf0b482 · inbound

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement cites this paper.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:49:01.387089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:8eb58de4205ea246d08eeed6db2df5bbdcb4df4a097d9d137c454e858dbe60e2