Pith. sign in

Paper Citation Record · LEDGER

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2507.21448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21448 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.951657Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.733134Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:25:47.176963Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d106859a-c6e3-4b29-85e9-e2332ce048a8 · outbound

This paper cites an unresolved cited work.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:50:08.154237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.727028Z digest=sha256:66bbae49f35a63da513a9c3f97995244ca722ff793a9b66dc4c5823962ef9d9c

Observation 28cb8503-5e9d-4c82-9740-8c72d3551bec · outbound

This paper cites Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.733134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.733134Z digest=sha256:96fa0be009eda65a44ff1cacd8086f264eb6b835ab9f0f1a03f757ffcbe4b2bd

Observation e9ee473f-0af7-45fd-a927-7e9eec61fc6a · outbound

This paper cites Datasets We use V oxCeleb2 as the speech dataset and MUSAN for noise and music.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Datasets We use V oxCeleb2 as the speech dataset and MUSAN for noise and music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:07.885831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.738101Z digest=sha256:858714c8e60b6203b38472765d89bc20c800e0eb2822118436fa195a8ddaace0

Observation fe688ff9-3a4b-415d-ad9d-7da2dd460b0b · outbound

This paper cites an unresolved cited work.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:50:07.587783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.743249Z digest=sha256:46de8421a5594584b2e700c1768392d4204cb465aa9eb1f70b4ad86e0e9f449e

Observation 043a9f23-38cb-400b-b8aa-b8a7632bcb31 · outbound

This paper cites We systemat- ically analyze how visual embeddings from audio-visual speech recognition (A VSR) and active speaker detection (ASD) im- pact A VSE performance.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations We systemat- ically analyze how visual embeddings from audio-visual speech recognition (A VSR) and active speaker detection (ASD) im- pact A VSE performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:07.283456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.748919Z digest=sha256:ae4b4127ab727c4d5028a365b4a45d1d4a07f8cf88f11186ab59f89841a13041

Observation 8cc567b3-95f6-49d6-affb-f61ab3fbe687 · outbound

This paper cites Funda- mentals, present and future perspectives of speech enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Funda- mentals, present and future perspectives of speech enhancement,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.957541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.754492Z digest=sha256:95fa4cf62c2e7bf4d88e5e33706416d13a60f60cd9adba9dfa1744b306bfe78c

Observation 6603e162-ab9a-48c5-87de-ad0fe3314878 · outbound

This paper cites An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.644010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.759259Z digest=sha256:5296ea972f0a3b5127c14ea1b2e81d724b5a1a03b7065a83b35312887294a72a

Observation a2b8f46e-57a4-4bc8-b7c1-2d3e777d5d76 · outbound

This paper cites FlowA VSE: Effi- cient Audio-Visual Speech Enhancement with Conditional Flow Matching,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations FlowA VSE: Effi- cient Audio-Visual Speech Enhancement with Conditional Flow Matching,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.324806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.764097Z digest=sha256:9625fae284952a67aab4206531887816dd0da4900df563d9a1e185392e6c505f

Observation f16c5e63-8812-4029-b90c-e8327b680cd1 · outbound

This paper cites Personalized speech enhancement: new models and Comprehensive evaluation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Personalized speech enhancement: new models and Comprehensive evaluation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.109228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.770003Z digest=sha256:408bb95a5c6db29f6928d2edee304c893e5878807feb389f9f821f2b7e62f3f7

Observation ec41fe69-5a12-4ac1-9fa3-8f61a6135da1 · outbound

This paper cites Real-Time Audio-Visual End-to-End Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual End-to-End Speech Enhancement,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.905454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.775199Z digest=sha256:e837de30627e2c426772be7586c74d40e33748319ece56bfcf83fc6afcdbf682

Observation 93e19642-7f0d-45c4-b5c1-82a5c30dd41d · outbound

This paper cites Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.752116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.780437Z digest=sha256:87ce63c20089a453a2a3336c42fdd403f444adb9bafdec5fe45834f9eb26b089

Observation 41dcc2f1-e283-4f1b-bfaf-dba2194b384e · outbound

This paper cites The Conversation: Deep Audio-Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations The Conversation: Deep Audio-Visual Speech Enhancement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.546179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.785047Z digest=sha256:a22ebd272643486f9d8654944b72f824373474a14f0f4be8e82d9110fbbc78f4

Observation 71d3d07a-a00f-4cdc-ad8b-52e049c05e8c · outbound

This paper cites Improved Lite Audio- Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Improved Lite Audio- Visual Speech Enhancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.219113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.790878Z digest=sha256:1bccb78cb8e0bd7050f6ceacd6082bb34ef24d81e85e2482a384427b9e9d0cf0

Observation 6d7ade20-feb2-4ca5-a041-e58855dbae5b · outbound

This paper cites Lite Audio- Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Lite Audio- Visual Speech Enhancement,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.981991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.800367Z digest=sha256:4922ba777f7919f8c9b43d35b13f1564dd3fe7253c5acb0a1ea7db0de6bde16f

Observation 36ccab49-88be-4571-84f6-0b66294fdd9f · outbound

This paper cites End-to-end Audio-visual Speech Recognition with Conformers,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations End-to-end Audio-visual Speech Recognition with Conformers,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.721415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.806913Z digest=sha256:920cbbbc98cde93937e8906b6d14edf017bb8565a22c115cdba5c13da156285c

Observation ff992ea9-cec4-4d7a-8524-b16de94de606 · outbound

This paper cites Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.471416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.815173Z digest=sha256:7752d8d09aed0e00ce267ddc5f6f89270fc17bb09c5481fd09b78276a9be4021

Observation 58298bd1-13a6-4b4c-8c64-1303f715c0a7 · outbound

This paper cites A Novel Real-Time, Lightweight Chaotic-Encryption Scheme for Next- Generation Audio-Visual Hearing Aids,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations A Novel Real-Time, Lightweight Chaotic-Encryption Scheme for Next- Generation Audio-Visual Hearing Aids,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.286004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.819911Z digest=sha256:549d64d66608e568423c428035e4bca11d63a2455daa44ba87c3cfd48810b33a

Observation bfb62780-2c0f-496b-9138-d95029e2e822 · outbound

This paper cites Lip- Reading Driven Deep Learning Approach for Speech Enhance- ment,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Lip- Reading Driven Deep Learning Approach for Speech Enhance- ment,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.040336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.825003Z digest=sha256:d47cb07ad62bae2b09f4fc8c306610c10568d8e4dfeb630ebeb3f3c5db465010

Observation 456aca05-6cfe-4c8c-b3f7-9887d564587c · outbound

This paper cites Audio-visual speech enhancement using deep neural networks,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-visual speech enhancement using deep neural networks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.809284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.830370Z digest=sha256:d80dfa3bdd8cb6427056d9913a40bb7b700ec244a59418058dd20b5c3d41f3f5

Observation efe08a43-f7f2-46c9-977a-4fb752add33f · outbound

This paper cites Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.656679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.835131Z digest=sha256:3e7e5b7b1adf5e686668a4822b80d16b2d60bb81a7bb85a7c8fd6b364de47636

Observation a7608f3b-cb83-422c-a4a2-ef088fc628f1 · outbound

This paper cites Evaluating Audiovisual Source Separation in the Context of Video Conferencing,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Evaluating Audiovisual Source Separation in the Context of Video Conferencing,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.495437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.840466Z digest=sha256:4e1f6b575d61b73e31afb8d8759836c450211fdf88cce3deaa327e5611c2ad31

Observation bb2528b1-7047-4d06-95f7-4707018a963d · outbound

This paper cites CochleaNet: A robust language-independent audio-visual model for real-time speech enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations CochleaNet: A robust language-independent audio-visual model for real-time speech enhancement,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.283326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.844890Z digest=sha256:e860eb2021caad1ee6d8361a2b6112350e603c178b16603f1dc7f9c0feea187b

Observation dbbd2a9a-9b4c-4af3-b48e-b3414632f914 · outbound

This paper cites RT-LA-V ocE: Real- Time Low-SNR Audio-Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations RT-LA-V ocE: Real- Time Low-SNR Audio-Visual Speech Enhancement,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.110631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.850747Z digest=sha256:2c69716c295cb951de6110315bd0585ace6dd0bf77c5e98fa9bd38820df7f7dd

Observation f9666af7-49c0-4b77-9739-43f790274900 · outbound

This paper cites Exploring Tradeoffs in Models for Low-Latency Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Exploring Tradeoffs in Models for Low-Latency Speech Enhancement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.900687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.855736Z digest=sha256:2eee6391a2d7dbfc37fd8f3f5342ceb5802eeef67ed217fc848f6e1e7a9cd356

Observation 461f0d2d-4cb5-4788-ab00-c1ca4ef1c00a · outbound

This paper cites Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.659182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.860441Z digest=sha256:86567c64cd498149bbe9b8e079325cefce54a7b3d956efc9c0d9596f21dfbafd

Observation b1b86e74-3ec7-4364-90d5-71c309b90f24 · outbound

This paper cites On The Compensation Between Magnitude and Phase in Speech Separation.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations On The Compensation Between Magnitude and Phase in Speech Separation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:50:01.119128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.865694Z digest=sha256:18e35f887444d232ddcc0a4b36412a60184c8541cfd740f4b658e00afa538ed8

Observation 35634989-c821-4f5d-9702-041a675255ac · outbound

This paper cites Learn- ing Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Learn- ing Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.452880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.872092Z digest=sha256:586dde1e88298f78302a01439fd660e8a8c88143814f974342b4ccce59356ebc

Observation 4be04175-9aba-4d25-8bb7-4c1f9d31a392 · outbound

This paper cites Robust Self-Supervised Audio-Visual Speech Recognition,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Robust Self-Supervised Audio-Visual Speech Recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.254963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.877218Z digest=sha256:5c260e33bdd1d7f5db9af026baf427dc51102e8e82cff50f124472aaf589d37e

Observation c83f124f-531d-47be-bf7d-6ed43c7bc8ae · outbound

This paper cites Visual Speech Recognition for Multiple Languages in the Wild.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Visual Speech Recognition for Multiple Languages in the Wild

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:50:01.052996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.882251Z digest=sha256:851b107996de6ec79fc37125d7039d2d16e11186f3bb708f36354f0f0622abf5

Observation a4299ee8-9644-4ef2-b76c-783a57c86573 · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.887904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.887904Z digest=sha256:b9f561cd2add22cdbe04cf2380d964c128922e78da2b4cd0060b983c3f078ad8

Observation 6e8efe7a-2f6a-453d-bc3b-1a8a4c1d2a22 · outbound

This paper cites LoCoNet: Long-Short Context Network for Active Speaker Detection,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations LoCoNet: Long-Short Context Network for Active Speaker Detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.049957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.894029Z digest=sha256:0fc149ade1dced04b9225b2a7ef8b45ac3bffbccde28fdc5276a687b00640e06

Observation 623afb87-27a1-420f-b9e9-92d3d3dd2098 · outbound

This paper cites Deep Contextualized Word Representa- tions,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Deep Contextualized Word Representa- tions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.878281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.899016Z digest=sha256:92f03c7489f3e1b39dd21885e3aa8ee860b8ca38c9164e5bbaf4e715b0af173f

Observation 7ad79254-0c6a-4a76-add1-0af98ccf5185 · outbound

This paper cites Few- shot Image Classification: Just Use a Library of Pre-trained Fea- ture Extractors and a Simple Classifier,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Few- shot Image Classification: Just Use a Library of Pre-trained Fea- ture Extractors and a Simple Classifier,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.651568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.904678Z digest=sha256:9ad20fdd1f718dcd1d38f4a5b8a994495f9179d429a6ad89e843502b378c3a9b

Observation 6cdf9559-167c-4aa0-9035-ae8f6ba41f81 · outbound

This paper cites Music auto-tagging in the long tail: A few-shot approach,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Music auto-tagging in the long tail: A few-shot approach,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.472179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.909319Z digest=sha256:22a970d093d186ed8b12ce889de1d4586ec2a356a8d7486c0a7ebe0dc7e34c12

Observation 7521eefd-a268-4a34-a654-eadff6bd0025 · outbound

This paper cites Contextual String Em- beddings for Sequence Labeling,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Contextual String Em- beddings for Sequence Labeling,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.398614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.918147Z digest=sha256:a9ca0d3ffe503e397827564f6594dfc9254bcd672e1093009480a726df95427b

Observation 736da3cf-bdca-444d-8d34-8482e8162dea · outbound

This paper cites DeepFilterNet: A Low Complexity Speech Enhancement Frame- work for Full-Band Audio based on Deep Filtering,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations DeepFilterNet: A Low Complexity Speech Enhancement Frame- work for Full-Band Audio based on Deep Filtering,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.362977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.922916Z digest=sha256:3fa9720bf11d881253588b166187655c568b8176bcbd3f1b89770a61b3633112

Observation 9a77b55c-e7fc-4f21-b061-7bf99b65d384 · outbound

This paper cites Perceptual eval- uation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Perceptual eval- uation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.319553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.929172Z digest=sha256:d05618987cd6c769f7e23259c6d33e3287eb37f24eb2adb1c7b1ea774dd0c7d6

Observation 791b7492-0965-4f57-8f69-f1fd60cf026d · outbound

This paper cites SDR – Half-baked or Well Done?.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations SDR – Half-baked or Well Done?

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.278688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.933900Z digest=sha256:afcd826900b4bc7d69bf570327b162a85eba9b429c8e437052a7f1b8224189ed

Observation cec44eee-4d81-4e17-a0fe-17ad826ce9d1 · outbound

This paper cites An Algorithm for Predicting the In- telligibility of Speech Masked by Modulated Noise Maskers,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations An Algorithm for Predicting the In- telligibility of Speech Masked by Modulated Noise Maskers,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.938902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.938902Z digest=sha256:3f8052444ed9dbc55faf71accb44f2087cc98c6e96102d846257a72f50b68baf

Observation 587209d4-2447-4d2d-ab72-2a0335df526d · outbound

This paper cites ViSpeR: Multilingual Audio-Visual Speech Recognition,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations ViSpeR: Multilingual Audio-Visual Speech Recognition,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.224069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.945904Z digest=sha256:b16551ddd924299c479da747cdc63ad4e820a91005ad55a63871216d2f29d434

Observation 4c3a8019-162a-48d6-b28a-553cc93946d2 · outbound

This paper cites MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.187419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:50:00.951657Z digest=sha256:adee7773f706583c0d64bdb53b6e48200b0322eea5a182acd25fce2c87fa5347

Pith citing papers

Observation 28cb8503-5e9d-4c82-9740-8c72d3551bec · inbound

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations cites this paper.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.733134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.733134Z digest=sha256:96fa0be009eda65a44ff1cacd8086f264eb6b835ab9f0f1a03f757ffcbe4b2bd

Observation b7317e6e-dbb0-42c4-8761-97e6f89fb4d3 · inbound

FSD50K-Solo: Automated Curation of Single-Source Sound Events cites this paper.

FSD50K-Solo: Automated Curation of Single-Source Sound Events Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:47.179446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T21:18:08.991522Z digest=sha256:5385a770188706e051535d13db81b42bad42253571dfb4e9d2307c623962093a