Pith. sign in

Paper Citation Record · LEDGER

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2507.21448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21448 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.951657Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:50:00.733134Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:25:47.176963Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d106859a-c6e3-4b29-85e9-e2332ce048a8 · outbound

This paper cites an unresolved cited work.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:50:08.154237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.727028Z digest=sha256:aaad6f1328989db61544493ec665eba82eca7e6db9e4938c139f63c96a6e2541

Observation 28cb8503-5e9d-4c82-9740-8c72d3551bec · outbound

This paper cites Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.733134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.733134Z digest=sha256:96fa0be009eda65a44ff1cacd8086f264eb6b835ab9f0f1a03f757ffcbe4b2bd

Observation e9ee473f-0af7-45fd-a927-7e9eec61fc6a · outbound

This paper cites Datasets We use V oxCeleb2 as the speech dataset and MUSAN for noise and music.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Datasets We use V oxCeleb2 as the speech dataset and MUSAN for noise and music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:07.885831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.738101Z digest=sha256:f82edd5083db81457c2a7d6d57554fcde505ac4a3f4f14c2d6c9be97773d97aa

Observation fe688ff9-3a4b-415d-ad9d-7da2dd460b0b · outbound

This paper cites an unresolved cited work.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:50:07.587783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.743249Z digest=sha256:5f0e816be98c387a251d65e82bff2fce6e5b9fe9625fbd48e8b2b923dc979ac4

Observation 043a9f23-38cb-400b-b8aa-b8a7632bcb31 · outbound

This paper cites We systemat- ically analyze how visual embeddings from audio-visual speech recognition (A VSR) and active speaker detection (ASD) im- pact A VSE performance.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations We systemat- ically analyze how visual embeddings from audio-visual speech recognition (A VSR) and active speaker detection (ASD) im- pact A VSE performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:07.283456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.748919Z digest=sha256:e27ceaebd37f37c69dbd4b1da4d90bbeca0625bd7a2088c44b734c45d5bb602f

Observation 8cc567b3-95f6-49d6-affb-f61ab3fbe687 · outbound

This paper cites Funda- mentals, present and future perspectives of speech enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Funda- mentals, present and future perspectives of speech enhancement,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.957541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.754492Z digest=sha256:110c675d0285b00c136111adb852242dfede757469806775fea7513e901cd3db

Observation 6603e162-ab9a-48c5-87de-ad0fe3314878 · outbound

This paper cites An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.644010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.759259Z digest=sha256:ffef2be3d5d89e4e141d4f55c5638b6c2fa7887383099d9c5ee6541d59e6b307

Observation a2b8f46e-57a4-4bc8-b7c1-2d3e777d5d76 · outbound

This paper cites FlowA VSE: Effi- cient Audio-Visual Speech Enhancement with Conditional Flow Matching,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations FlowA VSE: Effi- cient Audio-Visual Speech Enhancement with Conditional Flow Matching,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.324806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.764097Z digest=sha256:ac4046663816cc3e19a873dcc76c7cc9e9569ec087abe28c3660f6a085dad3d1

Observation f16c5e63-8812-4029-b90c-e8327b680cd1 · outbound

This paper cites Personalized speech enhancement: new models and Comprehensive evaluation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Personalized speech enhancement: new models and Comprehensive evaluation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:06.109228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.770003Z digest=sha256:0e2f9e1d237167abe257c32790e0561901718b730e9a511952ac4241be88b513

Observation ec41fe69-5a12-4ac1-9fa3-8f61a6135da1 · outbound

This paper cites Real-Time Audio-Visual End-to-End Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual End-to-End Speech Enhancement,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.905454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.775199Z digest=sha256:841fb42b30c5613972913114e7662ddfc87746fa3e4659e8ac388922513a72c6

Observation 93e19642-7f0d-45c4-b5c1-82a5c30dd41d · outbound

This paper cites Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.752116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.780437Z digest=sha256:cb03695e40539a242bee6ccc7fc86af407043404ac57742228240e6eb9a84797

Observation 41dcc2f1-e283-4f1b-bfaf-dba2194b384e · outbound

This paper cites The Conversation: Deep Audio-Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations The Conversation: Deep Audio-Visual Speech Enhancement,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.546179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.785047Z digest=sha256:8d7423006de65a1beb8dba7379dcdfed84d39248f959aac2ae36477b8afddee8

Observation 71d3d07a-a00f-4cdc-ad8b-52e049c05e8c · outbound

This paper cites Improved Lite Audio- Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Improved Lite Audio- Visual Speech Enhancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:05.219113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.790878Z digest=sha256:53d09cb7e9db0b889d001bed08b3e184a9c03f983ecf8dd96c06ad1a09f84598

Observation 6d7ade20-feb2-4ca5-a041-e58855dbae5b · outbound

This paper cites Lite Audio- Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Lite Audio- Visual Speech Enhancement,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.981991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.800367Z digest=sha256:9ed8978bef686d25a3b9cf72dcb2a902a475ecd6b4d18ee42a1479c48e0db5ee

Observation 36ccab49-88be-4571-84f6-0b66294fdd9f · outbound

This paper cites End-to-end Audio-visual Speech Recognition with Conformers,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations End-to-end Audio-visual Speech Recognition with Conformers,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.721415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.806913Z digest=sha256:3fea1d3b0271ccd4f87f721af99a313b758d3e0c7483b7f6db76b5c5129e75c9

Observation ff992ea9-cec4-4d7a-8524-b16de94de606 · outbound

This paper cites Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.471416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.815173Z digest=sha256:29307d8fb96801958e2ed62b8eab4113efdee26afb94ab5e5497cb0681977335

Observation 58298bd1-13a6-4b4c-8c64-1303f715c0a7 · outbound

This paper cites A Novel Real-Time, Lightweight Chaotic-Encryption Scheme for Next- Generation Audio-Visual Hearing Aids,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations A Novel Real-Time, Lightweight Chaotic-Encryption Scheme for Next- Generation Audio-Visual Hearing Aids,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.286004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.819911Z digest=sha256:21f415824c9c8070ef22c8cf4b8a5019a8bb7e29d982a4177c9b18c7936bd92e

Observation bfb62780-2c0f-496b-9138-d95029e2e822 · outbound

This paper cites Lip- Reading Driven Deep Learning Approach for Speech Enhance- ment,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Lip- Reading Driven Deep Learning Approach for Speech Enhance- ment,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:04.040336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.825003Z digest=sha256:a5250717f76600175ec45603da420baa24fbba812668b4299e4104d6178f1b2d

Observation 456aca05-6cfe-4c8c-b3f7-9887d564587c · outbound

This paper cites Audio-visual speech enhancement using deep neural networks,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-visual speech enhancement using deep neural networks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.809284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.830370Z digest=sha256:ed36806e2c23011fdd997da201eae413a870d932b891191b37579f401a575d9d

Observation efe08a43-f7f2-46c9-977a-4fb752add33f · outbound

This paper cites Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.656679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.835131Z digest=sha256:58966c265e3b9ff74f1e0ac284e6578ef4259e31ad479fdb3c4df70a54a17fe4

Observation a7608f3b-cb83-422c-a4a2-ef088fc628f1 · outbound

This paper cites Evaluating Audiovisual Source Separation in the Context of Video Conferencing,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Evaluating Audiovisual Source Separation in the Context of Video Conferencing,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.495437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.840466Z digest=sha256:51fb1321f55b1128c896d1926697c03049f9a6ebcf6a311881eac476a34c7ec1

Observation bb2528b1-7047-4d06-95f7-4707018a963d · outbound

This paper cites CochleaNet: A robust language-independent audio-visual model for real-time speech enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations CochleaNet: A robust language-independent audio-visual model for real-time speech enhancement,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.283326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.844890Z digest=sha256:90e6a15b8c48deec2d0d78eab2de6cb8a7ac48f28e78bda202a3c2b4620d401d

Observation dbbd2a9a-9b4c-4af3-b48e-b3414632f914 · outbound

This paper cites RT-LA-V ocE: Real- Time Low-SNR Audio-Visual Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations RT-LA-V ocE: Real- Time Low-SNR Audio-Visual Speech Enhancement,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:03.110631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.850747Z digest=sha256:4a1e8edddb8309781077ea1325971b9e2adb172aeb4e9c57910acd248ccfb9a7

Observation f9666af7-49c0-4b77-9739-43f790274900 · outbound

This paper cites Exploring Tradeoffs in Models for Low-Latency Speech Enhancement,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Exploring Tradeoffs in Models for Low-Latency Speech Enhancement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.900687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.855736Z digest=sha256:44249455ffc6f137915bb2efd072d7cddb10b0c26edfd4bd088320e8c93484f5

Observation 461f0d2d-4cb5-4788-ab00-c1ca4ef1c00a · outbound

This paper cites Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.659182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.860441Z digest=sha256:1178a3693eca939c189344e9a7e3005fd5dab224850e8110b1fb0badb054b3ac

Observation b1b86e74-3ec7-4364-90d5-71c309b90f24 · outbound

This paper cites On The Compensation Between Magnitude and Phase in Speech Separation.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations On The Compensation Between Magnitude and Phase in Speech Separation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:50:01.119128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.865694Z digest=sha256:a2ccb6474fd9ff9c564a6c7af11e8a8859ca4b578b1f0744ee769f04a0fb0695

Observation 35634989-c821-4f5d-9702-041a675255ac · outbound

This paper cites Learn- ing Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Learn- ing Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.452880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.872092Z digest=sha256:04f6d6de481c87f3ce500f76a96301b536e7e29f17f8ece0934305d2ab7db61b

Observation 4be04175-9aba-4d25-8bb7-4c1f9d31a392 · outbound

This paper cites Robust Self-Supervised Audio-Visual Speech Recognition,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Robust Self-Supervised Audio-Visual Speech Recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.254963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.877218Z digest=sha256:42c1252782efe3b5c081e76d924387fe59489b0f35f08e48be3553ad29c43dcb

Observation c83f124f-531d-47be-bf7d-6ed43c7bc8ae · outbound

This paper cites Visual Speech Recognition for Multiple Languages in the Wild.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Visual Speech Recognition for Multiple Languages in the Wild

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:50:01.052996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.882251Z digest=sha256:58b7822972d791c40ca217a75d41b69fca623d0a3e7c7605959640a9e4189079

Observation a4299ee8-9644-4ef2-b76c-783a57c86573 · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.887904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.887904Z digest=sha256:251cb0338257fe3bfd54cadc998a6388754351f3a76a9106ed1d3d9a94e66608

Observation 6e8efe7a-2f6a-453d-bc3b-1a8a4c1d2a22 · outbound

This paper cites LoCoNet: Long-Short Context Network for Active Speaker Detection,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations LoCoNet: Long-Short Context Network for Active Speaker Detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:02.049957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.894029Z digest=sha256:b972bb4207fc80ba8c885a5da51411d4e5fcc5fa01d31e5084990bc09286146a

Observation 623afb87-27a1-420f-b9e9-92d3d3dd2098 · outbound

This paper cites Deep Contextualized Word Representa- tions,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Deep Contextualized Word Representa- tions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.878281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.899016Z digest=sha256:256a5ab6f6bfffe8175900d0ec8f7515d9298750cdcdea1970b40afa8ec5c0a7

Observation 7ad79254-0c6a-4a76-add1-0af98ccf5185 · outbound

This paper cites Few- shot Image Classification: Just Use a Library of Pre-trained Fea- ture Extractors and a Simple Classifier,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Few- shot Image Classification: Just Use a Library of Pre-trained Fea- ture Extractors and a Simple Classifier,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.651568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.904678Z digest=sha256:2d6989580da4733f6764a5188a24e12c7b8c654ef1375d1a7ea9f45f237c2c9a

Observation 6cdf9559-167c-4aa0-9035-ae8f6ba41f81 · outbound

This paper cites Music auto-tagging in the long tail: A few-shot approach,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Music auto-tagging in the long tail: A few-shot approach,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.472179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.909319Z digest=sha256:96e0604752366b36909e3f836e1d4777d5c21827655eb86a92cc395c915811dd

Observation 7521eefd-a268-4a34-a654-eadff6bd0025 · outbound

This paper cites Contextual String Em- beddings for Sequence Labeling,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Contextual String Em- beddings for Sequence Labeling,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.398614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.918147Z digest=sha256:c1e31009da93b4074a6104ae75b94721876c71413c0ff30dd0f0c03792ea2092

Observation 736da3cf-bdca-444d-8d34-8482e8162dea · outbound

This paper cites DeepFilterNet: A Low Complexity Speech Enhancement Frame- work for Full-Band Audio based on Deep Filtering,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations DeepFilterNet: A Low Complexity Speech Enhancement Frame- work for Full-Band Audio based on Deep Filtering,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.362977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.922916Z digest=sha256:1d52c1ac0a0719f59cf372e1bfc72048d92aa0c3ecc619eaa0649682ec405b23

Observation 9a77b55c-e7fc-4f21-b061-7bf99b65d384 · outbound

This paper cites Perceptual eval- uation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Perceptual eval- uation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.319553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.929172Z digest=sha256:ea7889fab9eb28f465f64c4412705a949bbdaab4e2ac8cdd4cdd0c7d63288f76

Observation 791b7492-0965-4f57-8f69-f1fd60cf026d · outbound

This paper cites SDR – Half-baked or Well Done?.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations SDR – Half-baked or Well Done?

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.278688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.933900Z digest=sha256:36175352cbbb5cd9101c6b23cfa441ddd2b2e0b60c63140f2fdd1cc798891ef7

Observation cec44eee-4d81-4e17-a0fe-17ad826ce9d1 · outbound

This paper cites An Algorithm for Predicting the In- telligibility of Speech Masked by Modulated Noise Maskers,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations An Algorithm for Predicting the In- telligibility of Speech Masked by Modulated Noise Maskers,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.938902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.938902Z digest=sha256:3f8052444ed9dbc55faf71accb44f2087cc98c6e96102d846257a72f50b68baf

Observation 587209d4-2447-4d2d-ab72-2a0335df526d · outbound

This paper cites ViSpeR: Multilingual Audio-Visual Speech Recognition,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations ViSpeR: Multilingual Audio-Visual Speech Recognition,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.224069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.945904Z digest=sha256:b0fc6e40a804337b10663f22f7e3194561dd2a673de21feedc66733dd682c6dc

Observation 4c3a8019-162a-48d6-b28a-553cc93946d2 · outbound

This paper cites MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation,.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations MEAD: A Large-Scale Audio-Visual Dataset for Emotional Talking-Face Generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:50:01.187419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:50:00.951657Z digest=sha256:03ae74304baee105aebeaf0b14277313d1e6a70e7565b4434783e8cfdb399947

Pith citing papers

Observation 28cb8503-5e9d-4c82-9740-8c72d3551bec · inbound

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations cites this paper.

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:50:00.733134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:50:00.733134Z digest=sha256:96fa0be009eda65a44ff1cacd8086f264eb6b835ab9f0f1a03f757ffcbe4b2bd

Observation b7317e6e-dbb0-42c4-8761-97e6f89fb4d3 · inbound

FSD50K-Solo: Automated Curation of Single-Source Sound Events cites this paper.

FSD50K-Solo: Automated Curation of Single-Source Sound Events Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:47.179446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:18:08.991522Z digest=sha256:97cbc95cb4e5766834b8f66d624a073f749cf0224335fbeb2dfad556830b964d