Pith. sign in

Paper Citation Record · LEDGER

Online Audio-Visual Autoregressive Speaker Extraction

As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2506.01270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01270 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:49.884180Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:45.946937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:50:50.003069Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0bb2916-2367-4544-8184-3e624188d338 · outbound

This paper cites Online Audio-Visual Autoregressive Speaker Extraction.

Online Audio-Visual Autoregressive Speaker Extraction Online Audio-Visual Autoregressive Speaker Extraction

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:50:50.031211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:45.946937Z digest=sha256:ece3fbd62c98dc34c6f45ccb9656a73a3e2d55d3d954d6feee979e272eeafdf3

Observation ce03a151-9cda-4578-93e7-be345624fcf8 · outbound

This paper cites pseudo past extracted speech.

Online Audio-Visual Autoregressive Speaker Extraction pseudo past extracted speech

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.601221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.061921Z digest=sha256:99637538b75929cb84b5c59ce8c25ed1a5a13b2afb85b69c07494eb3904a05cd

Observation ec701100-3023-4995-92f5-82ac578580bf · outbound

This paper cites Dataset We mainly use the Lip Reading Sentences 3 (LRS3) dataset to validate our proposed method in this work [29], which is widely used in many A VSE studies [34–36].

Online Audio-Visual Autoregressive Speaker Extraction Dataset We mainly use the Lip Reading Sentences 3 (LRS3) dataset to validate our proposed method in this work [29], which is widely used in many A VSE studies [34–36]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.548076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.143478Z digest=sha256:df08b36e359cca05acc63f263e2bdb72dcc159bb40895ce99221a7058b707ee7

Observation c4fd1eef-0518-4b67-bae9-21b283447642 · outbound

This paper cites All improve- ments are calculated relative to the unprocessed multi-talker speech signals.

Online Audio-Visual Autoregressive Speaker Extraction All improve- ments are calculated relative to the unprocessed multi-talker speech signals

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:50:54.505498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.230175Z digest=sha256:d052d35b3bbba71d6e13356c736faabb70f300a2a2b88a4875201d1441545a95

Observation e69b16ad-948b-4e99-82a6-270fe3813ccc · outbound

This paper cites The proposed visual en- coder, with its lightweight design and efficient processing, pro- vides a competitive alternative to the more complex visual en- coder.

Online Audio-Visual Autoregressive Speaker Extraction The proposed visual en- coder, with its lightweight design and efficient processing, pro- vides a competitive alternative to the more complex visual en- coder

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.444857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.312156Z digest=sha256:f95d3c4849a7be85892f0cf63b842130a13cadcfc729ea28b4bd42bc11638f67

Observation 4c98a724-67a7-4f3f-bb5c-f388384a8ce9 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

Online Audio-Visual Autoregressive Speaker Extraction Some experiments on the recognition of speech, with one and with two ears,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.370756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.447958Z digest=sha256:61dd1ce2bd869bca7cb4137801204a2922f8f032548adb5d4cf41b465cd6e541

Observation acacf362-dc1b-416a-8cd0-0c191d56ce1e · outbound

This paper cites Restoring speaking lips from occlusion for audio-visual speech recognition,.

Online Audio-Visual Autoregressive Speaker Extraction Restoring speaking lips from occlusion for audio-visual speech recognition,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.319219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.558333Z digest=sha256:739b494e482c1d3dd2ac086a210e189064359cb1a6b10c408978802a4ec9077c

Observation 8a23a9c5-99ca-4986-88e3-c1be43aaf497 · outbound

This paper cites Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,.

Online Audio-Visual Autoregressive Speaker Extraction Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.266693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.663185Z digest=sha256:5de9276da7015fdb306c10cd7fdbe26c7f83db339fd1a575a97f81885b312814

Observation 36529006-3bd3-44f2-8c80-b56bd9a14608 · outbound

This paper cites Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,.

Online Audio-Visual Autoregressive Speaker Extraction Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.200466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.762989Z digest=sha256:f70d188bf45f7c24287568eb457795088615af83434920bc7b0ee7899b9a0a29

Observation f9d9b155-52f7-42cc-a7ad-9b5ba50224aa · outbound

This paper cites Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

Online Audio-Visual Autoregressive Speaker Extraction Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:54.095360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:46.873189Z digest=sha256:444a9c548acd5798ec187b5270b961898d1b434bb0b5117c758522a57b318ec3

Observation d5aaa220-4d04-42d1-9e13-1cf09631eb0a · outbound

This paper cites TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,.

Online Audio-Visual Autoregressive Speaker Extraction TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.956566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.025238Z digest=sha256:3ae8164ddb878545b4852580de3ddda47fd132e9831bb72fc2309d6ddd6bc9d6

Observation da0bb5f4-376a-45d2-ae30-93b40f7e729a · outbound

This paper cites Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,.

Online Audio-Visual Autoregressive Speaker Extraction Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.871922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.108156Z digest=sha256:4a8cdfcccf0360da4752b32231edb492f510cb09b4c7d512d8a8de461e82d088

Observation 06a1a5a5-0872-4348-84f2-cde2b38d109f · outbound

This paper cites V oice- Filter: Targeted voice separation by speaker-conditioned spectro- gram masking,.

Online Audio-Visual Autoregressive Speaker Extraction V oice- Filter: Targeted voice separation by speaker-conditioned spectro- gram masking,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.775445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.185101Z digest=sha256:d549c4447a4230760ca267b2d1bb2e69ccdc0b0919c56e1544663f8dd52a84ce

Observation 0bcb3c14-7c28-4a3e-8859-7ed55aea51c1 · outbound

This paper cites SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

Online Audio-Visual Autoregressive Speaker Extraction SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.681268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.287943Z digest=sha256:5c80753a6e719d2c13e93f31ce2a2fe1acd435bedf4624eacb5a05c192e09d31

Observation d0b41310-99cd-49b6-99a6-1a1969c83399 · outbound

This paper cites SpEx: Multi-scale time domain speaker extraction network,.

Online Audio-Visual Autoregressive Speaker Extraction SpEx: Multi-scale time domain speaker extraction network,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.573479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.368732Z digest=sha256:174d418b004120b30800f1687a3b7d945a3b44e652c5f3934c9890ec09a07389

Observation 862ee3c5-15eb-4d2d-9b62-f452c8757511 · outbound

This paper cites Looking to listen at the cock- tail party: a speaker-independent audio-visual model for speech separation,.

Online Audio-Visual Autoregressive Speaker Extraction Looking to listen at the cock- tail party: a speaker-independent audio-visual model for speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.489372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.435865Z digest=sha256:422cc1088f423de72f86a87236427c99d58f41b000cac8d9bfb89892834b0927

Observation a9e535be-dc5c-4eec-8e24-6bd511b70bc3 · outbound

This paper cites Scenario-aware audio-visual TF- Gridnet for target speech extraction,.

Online Audio-Visual Autoregressive Speaker Extraction Scenario-aware audio-visual TF- Gridnet for target speech extraction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.391924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.540980Z digest=sha256:62523ed04cfdec56f4837f3139329b87b7756777ba80485193c305c0dc42f01b

Observation 4e460beb-d76d-48e7-b698-c1f08372a26f · outbound

This paper cites Time domain audio visual speech separation,.

Online Audio-Visual Autoregressive Speaker Extraction Time domain audio visual speech separation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.273195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.661462Z digest=sha256:f5006f4d90673d9b7c8c1583e33c510fb0a835f04466e1aa87830c1773b0d404

Observation 959f2c3a-c0eb-4694-89e5-c85f72243a22 · outbound

This paper cites Selective listening by synchro- nizing speech with lips,.

Online Audio-Visual Autoregressive Speaker Extraction Selective listening by synchro- nizing speech with lips,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.167144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.774934Z digest=sha256:12e2d63a8af049a81d222d4615b25b2712ff0be6b845b648c1276fbe967486d6

Observation 9f38ef15-19ca-418a-8d91-2e61a275589f · outbound

This paper cites FaceFilter: Audio-visual speech separation using still images,.

Online Audio-Visual Autoregressive Speaker Extraction FaceFilter: Audio-visual speech separation using still images,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:53.071792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:47.896449Z digest=sha256:f3332bda89afe2a3defcaed8da25fe8b616e54807a39c065709a4ced60d7b6c7

Observation 3c11cfa0-3fa1-482c-ae8a-a436d0d73c27 · outbound

This paper cites Brain- informed speech separation (BISS) for enhancement of target speaker in multitalker speech perception,.

Online Audio-Visual Autoregressive Speaker Extraction Brain- informed speech separation (BISS) for enhancement of target speaker in multitalker speech perception,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.968873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.010934Z digest=sha256:4e18fe705bfc4ab24311d3bbdd1d1d0110a79b87c725eaa0b06f2db442388375

Observation 5f239228-c3cf-4fd8-a929-251aae6a68f3 · outbound

This paper cites NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals.

Online Audio-Visual Autoregressive Speaker Extraction NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:48.079784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:48.079784Z digest=sha256:adc529ab32ff3dc0d97c1b5b4d268aaa00fb98e7e7d7d67fb90f9a1d3af2cda9

Observation 8e02594b-dcc2-4169-b27a-699abd8ab9cb · outbound

This paper cites NeuroHeed+: Improving neuro-steered speaker extraction with joint auditory attention detection,.

Online Audio-Visual Autoregressive Speaker Extraction NeuroHeed+: Improving neuro-steered speaker extraction with joint auditory attention detection,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.892960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.181836Z digest=sha256:abff7c50e0be4632b5fdfde9220a6e3f9cd856ca684a0afa63cdad38bed5d0b0

Observation b1beb070-f189-46f9-b306-c2d40b487bb3 · outbound

This paper cites V oiceFilter-Lite: Streaming targeted voice separation for on-device speech recog- nition,.

Online Audio-Visual Autoregressive Speaker Extraction V oiceFilter-Lite: Streaming targeted voice separation for on-device speech recog- nition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.825277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.300416Z digest=sha256:6ca7e7f703b447b0a9f0fd5e87e2a0a45e6780fa843e3e97c2c7c95dc7a8f0ca

Observation 866ccd12-e1fe-49ff-b9cc-47965431eab0 · outbound

This paper cites Papez: Resource-efficient speech sepa- ration with auditory working memory,.

Online Audio-Visual Autoregressive Speaker Extraction Papez: Resource-efficient speech sepa- ration with auditory working memory,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.737790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.411506Z digest=sha256:9f35e3b88abc1707a04581b1089b104915b207d3ecc950c5f8e65df8cf06c345

Observation 1fa3fe15-e872-41fb-9b66-c67efa9a29b5 · outbound

This paper cites SkiM: Skipping memory LSTM for low-latency real-time continuous speech separation,.

Online Audio-Visual Autoregressive Speaker Extraction SkiM: Skipping memory LSTM for low-latency real-time continuous speech separation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.619213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.553960Z digest=sha256:094799e053358016306466003ae2a363b5eed9526b811a4dc37feb2e3d24db81

Observation afbcb701-099d-4d62-b252-086dbd33e94a · outbound

This paper cites Resource-efficient separation transformer,.

Online Audio-Visual Autoregressive Speaker Extraction Resource-efficient separation transformer,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.513483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.687287Z digest=sha256:6589a740a41e219cf5f8abdf65488c530d92d5a29fb7284b81683be118709f32

Observation a2d34585-d1dd-4d1f-ace2-1381f1283e16 · outbound

This paper cites RT-LA-V ocE: Real- time low-SNR audio-visual speech enhancement,.

Online Audio-Visual Autoregressive Speaker Extraction RT-LA-V ocE: Real- time low-SNR audio-visual speech enhancement,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.414327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.768089Z digest=sha256:f6bda69d1ff43965dcb8e3cd1da01efca69d0c7d2138b2efe02825adc0a3477f

Observation 2fc7b33a-982c-4b07-871f-386bbaf5cf5e · outbound

This paper cites USEV: Universal speaker extraction with visual cue,.

Online Audio-Visual Autoregressive Speaker Extraction USEV: Universal speaker extraction with visual cue,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.298472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.873704Z digest=sha256:c97869664d08f9e2b812228b1d9bb5183cab786d95146aae936a9c962e9175bc

Observation 0fe65e12-aa9c-416e-a257-3586680a3dcf · outbound

This paper cites Real-time audio-visual end-to-end speech enhance- ment,.

Online Audio-Visual Autoregressive Speaker Extraction Real-time audio-visual end-to-end speech enhance- ment,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.205859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:48.969007Z digest=sha256:c2ef915ed4b1edbf16fe3bcee1d15d2a55e8a11ac0507f66eaa3fc570cb07048

Observation 8db1e0e9-c7b2-4c06-9935-569a61406b40 · outbound

This paper cites BlazeFace: Sub-millisecond neural face detec- tion on mobile GPUs,.

Online Audio-Visual Autoregressive Speaker Extraction BlazeFace: Sub-millisecond neural face detec- tion on mobile GPUs,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:52.058424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.032778Z digest=sha256:9edc694ec6099311f2954c53ab0eb5ff1b31321689fb1abf39aa34aa2f7d6334

Observation 1a6f1f38-727e-4965-9ebd-232eee34fb2c · outbound

This paper cites Xception: Deep learning with depthwise separable convolutions,.

Online Audio-Visual Autoregressive Speaker Extraction Xception: Deep learning with depthwise separable convolutions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.893803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.078134Z digest=sha256:ac49100e00f393c376b9824593eefd079f6819585fc3c59e3bccb15b9cc92bcf

Observation 8da9928e-d137-4ee7-9786-714afc877e7f · outbound

This paper cites PARIS: Pseudo-autoregressive siamese training for online speech separation,.

Online Audio-Visual Autoregressive Speaker Extraction PARIS: Pseudo-autoregressive siamese training for online speech separation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.789307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.123881Z digest=sha256:e5584ccba34e0b79e37d89f39caed6db189f09da2d644453b01f9fde24a9c80d

Observation e3be1cbd-ca63-4c40-8f67-cf50e492bb74 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Online Audio-Visual Autoregressive Speaker Extraction LRS3-TED: a large-scale dataset for visual speech recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:49.167617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:49.167617Z digest=sha256:4c3751b0ad296bcdddd0ff9c76bcc2cb6cb671c86805a13ee0dbc52a0dda8d81

Observation 57341b6c-1d97-4a12-9078-2ee72405c140 · outbound

This paper cites Deep lip reading: A comparison of models and an online application,.

Online Audio-Visual Autoregressive Speaker Extraction Deep lip reading: A comparison of models and an online application,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.705874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.212256Z digest=sha256:9ec7e5cedd2f04e90fa988a682dc58b7b8227b055abec3a22327fd0f2dd2f579

Observation ded28af7-8f67-41cd-adad-d2195784ee71 · outbound

This paper cites MuSE: Multi-modal target speaker extraction with visual cues,.

Online Audio-Visual Autoregressive Speaker Extraction MuSE: Multi-modal target speaker extraction with visual cues,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.504364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.264090Z digest=sha256:b1c58e1051dde2241b4ca91485a14f10e54e6fcb54f9cab9a49a6c4f7e50e0d6

Observation 993a4242-7a2b-4b70-8726-b71ff01b3839 · outbound

This paper cites SDR– half-baked or well done?.

Online Audio-Visual Autoregressive Speaker Extraction SDR– half-baked or well done?

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.325118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.313436Z digest=sha256:bc5963beec02e12fcabc97f9da3fe266ec0035bb3eeadf51f1aa433c3fc0f5ca

Observation 000c2997-b4e3-464e-95b6-6ce93a75e03c · outbound

This paper cites A hybrid continuity loss to reduce over-suppression for time-domain target speaker extraction,.

Online Audio-Visual Autoregressive Speaker Extraction A hybrid continuity loss to reduce over-suppression for time-domain target speaker extraction,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.141121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.370327Z digest=sha256:29bad73929d384fc450b9eef4001c7e59e4cfcff81d0b37d0491c0bf46621902

Observation 506f3ef7-487d-4a88-9952-9d0921026f86 · outbound

This paper cites ReVISE: Self-supervised speech resynthesis with visual input for universal and generalized speech regeneration,.

Online Audio-Visual Autoregressive Speaker Extraction ReVISE: Self-supervised speech resynthesis with visual input for universal and generalized speech regeneration,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:51.007488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.407524Z digest=sha256:dd2882e27bd67f1b2f1d419b54f9061c1286fc3d39ec5c1434f8a73c9006156d

Observation bfb8f8cd-2e02-4c74-ac72-95ab2da7280b · outbound

This paper cites PIA VE: A pose-invariant audio- visual speaker extraction network,.

Online Audio-Visual Autoregressive Speaker Extraction PIA VE: A pose-invariant audio- visual speaker extraction network,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:50.883044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.451366Z digest=sha256:876cea4654070d76ca679c9965e9921558d33cce0eb7f1ff90ff0f6db4866d90

Observation 6722909d-3a70-46f3-9a83-0ef8ada3c5d3 · outbound

This paper cites Audio- visual speech separation in noisy environments with a lightweight iterative model,.

Online Audio-Visual Autoregressive Speaker Extraction Audio- visual speech separation in noisy environments with a lightweight iterative model,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:50.756506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.498451Z digest=sha256:789520a1a85c586fd4f725bfc09cf36301fde45614853d4f9f42a725297d63ec

Observation 9f4a7fef-d652-4b69-9a11-f4f5e11d8f50 · outbound

This paper cites Adam, a method for stochastic optimiza- tion,.

Online Audio-Visual Autoregressive Speaker Extraction Adam, a method for stochastic optimiza- tion,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:50.626463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.600183Z digest=sha256:8b4bca38278122c7d8b4c093b864bc69db3908ee9fb69b9785e5764a37ecd63a

Observation ef72e77f-0bb0-4e40-8879-c998dffb32e6 · outbound

This paper cites Performance mea- surement in blind audio source separation,.

Online Audio-Visual Autoregressive Speaker Extraction Performance mea- surement in blind audio source separation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:49.679700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:49.679700Z digest=sha256:425ca72d21021c7f1fb48d04b22468641483237d2db05e76bc3992c416bc028f

Observation 9370e7b9-421e-42cd-9b3a-03339db643a6 · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Online Audio-Visual Autoregressive Speaker Extraction Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:50.505960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.721577Z digest=sha256:5d69dc0586c88713c9c8d72017fa7d3fb3a56ccb20cb01198a4525fae55e86ca

Observation eed5317f-9f84-4542-be5a-f322163b464a · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

Online Audio-Visual Autoregressive Speaker Extraction A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:49.782532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:49.782532Z digest=sha256:7ce7ce6033137dab45316844017e335371d5ae0282dcfc480ac2bace741464fc

Observation 83bd3914-87fb-43d7-a216-331d466c124d · outbound

This paper cites V oxCeleb2: Deep speaker recognition,.

Online Audio-Visual Autoregressive Speaker Extraction V oxCeleb2: Deep speaker recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:50.342860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.831817Z digest=sha256:55ec5e9f345c9c3de5c8bbc6f8d1e537207dfc3a42a8bc55700688396b23a781

Observation 2d9ecaa1-10c4-408e-ab9c-fa32ea251442 · outbound

This paper cites TCD-TIMIT: An audio-visual corpus of continuous speech,.

Online Audio-Visual Autoregressive Speaker Extraction TCD-TIMIT: An audio-visual corpus of continuous speech,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:50:50.169266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:49.884180Z digest=sha256:53a13cfe8b7c61251f44c322cb9d74bb2d0b71ad4f7de2b04a283e46f2fd1117

Pith citing papers

Observation c0bb2916-2367-4544-8184-3e624188d338 · inbound

Online Audio-Visual Autoregressive Speaker Extraction cites this paper.

Online Audio-Visual Autoregressive Speaker Extraction Online Audio-Visual Autoregressive Speaker Extraction

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:50:50.031211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:50:45.946937Z digest=sha256:ece3fbd62c98dc34c6f45ccb9656a73a3e2d55d3d954d6feee979e272eeafdf3