Pith. sign in

Paper Citation Record · LEDGER

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2505.03186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03186 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:14.805352Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ac9d1eb-96a8-49c5-b3b3-662a4e4126b6 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Robust speech recognition via large-scale weak supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.575052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.575052Z digest=sha256:cb55411311915c18e8a9fdcd21f2fd6680af33c88bb5e842f76ba2d016d681f0

Observation 626beae2-331e-40c3-8d07-f7e07c6f1686 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.581161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.581161Z digest=sha256:6e6744471f35285fab48d0f49eae7af8855cf41e9cf90318a8c676c8646e98cb

Observation bbf19573-56c2-41df-8141-8454e2d90440 · outbound

This paper cites Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.586465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.586465Z digest=sha256:d7d159e76dabce7feeae51d00a2c6e510ecae7893f12a6b75ca667a94a7b11d0

Observation 89ee59dd-884d-426b-9b54-0dce5284d186 · outbound

This paper cites Qwen2-audio technical report, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Qwen2-audio technical report, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.592073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.592073Z digest=sha256:fd5261b02c55f6f34f4492d9d3d9306988def9ddb14636e437dd16352973abb7

Observation b8ae44e0-4685-4a19-88e0-25d0f1ebae40 · outbound

This paper cites Assessment for automatic speech recognition: Ii.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Assessment for automatic speech recognition: Ii

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.597519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.597519Z digest=sha256:a8b48688dd91723f64f936ba50253517fdef63088f74775088085312c73d7602

Observation 18e863d1-c7b0-4231-a73e-1fa10e3e797f · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.495059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.602336Z digest=sha256:4035244847dbe0f904a751c366276636c9f1e88b1fe9c1f018d7c7d552dc0611

Observation 31658b4e-5454-4c42-a0c0-c56795cce0fb · outbound

This paper cites mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:02:14.909462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.608210Z digest=sha256:f5f317a27d7518355f3220013201051a3b4502bcd6e054ed5dcec5146ac5e4ed

Observation 69e8cff1-ce87-47ee-8f48-4bf620f51cb2 · outbound

This paper cites Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.480421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.613614Z digest=sha256:85eff0f38e46ad662a7057e438059b7246eb362f0aa2fe17a90e94f17c7194c7

Observation a5ea468d-4a1f-4257-a40e-c9b49796ebc2 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.618734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.618734Z digest=sha256:d6c5d7b080ad29e99c94aae18b7a5bd9f72725a4a540c61b8da3f310575464cf

Observation 99d9da77-5ffd-4467-b0f9-98f8463fde04 · outbound

This paper cites Unified speech recognition: A single model for auditory, visual, and audiovisual inputs, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Unified speech recognition: A single model for auditory, visual, and audiovisual inputs, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.464167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.623425Z digest=sha256:b3fc84fee1f9498381bdb153ffa7eea4d4abdcc54efadbac25a90b6d3ca45c3f

Observation 4ee4a0e4-4941-4724-95f9-13f75958055a · outbound

This paper cites Speech recognition models are strong lip- readers.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Speech recognition models are strong lip- readers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.449473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.627623Z digest=sha256:8a1396f0d69d1cfefc439e6dadb176bbd545a740ad2e01b9e26ec97e8c89ed98

Observation 14646b13-963a-4262-9c08-d723ac268e36 · outbound

This paper cites Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.631987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.631987Z digest=sha256:4decb15a1c51acb0179fd64d059ed3e532d10cdcb06212e4d253a0c63a0e5842

Observation 5ea2047c-6fb3-401f-8059-8d4a7f85e7d0 · outbound

This paper cites van de Ven, Nicholas Soures, and Dhireesha Kudithipudi.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization van de Ven, Nicholas Soures, and Dhireesha Kudithipudi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.434451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.636993Z digest=sha256:dbb16c02fcfc66efcac9320fcf8c94e95bfdf27a43c0d35ad435b2b601a40ad9

Observation 6d79bc66-5e5f-44d1-b6b7-8df976f3894d · outbound

This paper cites Lip reading sentences in the wild.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lip reading sentences in the wild

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.420078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.641960Z digest=sha256:0419999532a2b665072aaa2f8f25c54e84ed4172eb545627bdf99265f98c35f6

Observation 9f2bd975-770e-4681-b5fb-98a978667072 · outbound

This paper cites Maas: Multi-modal assignation for active speaker detection.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Maas: Multi-modal assignation for active speaker detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.405666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.646482Z digest=sha256:3dba46ee7bc9c4ef2ff6eda4b723b98bdcf5089e5aa0708316ba53c8ed6b60cb

Observation 4e685917-7581-4da2-b044-28ecc04cadc3 · outbound

This paper cites Out of time: automated lip sync in the wild.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Out of time: automated lip sync in the wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.650794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.650794Z digest=sha256:916c3206055fd5266aa0708b5606febdc7993ec2e9fbc004405d365e8fb9de28

Observation 40a0b99c-cd91-4b60-a5a3-64a0b827c765 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization A lip sync expert is all you need for speech to lip generation in the wild

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.655093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.655093Z digest=sha256:7e53e14e7fa1fc4c506ab7b4f456174ad8853d24a97076f3c594d2f877bf907f

Observation 9faab5e2-a5a3-4722-84e2-dac9fc254347 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.659428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.659428Z digest=sha256:bd4146cc97918dbafbcf8afeb07da5d1b4544668bad646d57875ff4da055a9ef

Observation ae52f234-c1ae-4b5f-a74d-31bc48f3a648 · outbound

This paper cites Asr is all you need: cross-modal distillation for lip reading, 2020.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Asr is all you need: cross-modal distillation for lip reading, 2020

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.371184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.663974Z digest=sha256:7d4bca274fcb0a469c3b714e50fc6b1721f42354f9144de291482451fd78aae3

Observation 1108b9bb-53f2-4992-a221-18a83f7e9e60 · outbound

This paper cites Schuller, and Maja Pantic.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Schuller, and Maja Pantic

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.356605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.668210Z digest=sha256:649f20c8561cacb7151d9dcd9a1e3afa7e5ada4a22f3ce7e3ffcf4299063a53b

Observation 92fc2400-5d70-4ad9-91dc-86d271b22ca8 · outbound

This paper cites u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality, 2022.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.342794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.672686Z digest=sha256:d27e147f43eb4fcbf568bbeab14978647c8299270f6a98b87b19deadb818dea2

Observation 56b24ef1-c4cc-49f2-b99a-859acfff1e0a · outbound

This paper cites Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.328435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.676961Z digest=sha256:97a226add85a54f19b83e052a080017c6fd91579f8e3f4861cffc3bba5a32687

Observation 9ea9bd24-eaa7-4d29-802c-09159655c654 · outbound

This paper cites Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.314319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.681001Z digest=sha256:050dc7828846c3cc94b556a05e0a28c5016dcaf2fe296452bacd8ebf0e561546

Observation 38f8b3c4-a53a-4c7c-bd27-8c4eca718d53 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition, 2020.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Conformer: Convolution-augmented transformer for speech recognition, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.685115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.685115Z digest=sha256:359f79cf1b6861d49acbc91c63ca4b6e70bff89aff49d53ed614cb3e71e17ad5

Observation 87103140-2d4c-47aa-a9fa-5dfd522edce5 · outbound

This paper cites Multilingual audio-visual speech recognition with hybrid ctc/rnn-t fast conformer.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Multilingual audio-visual speech recognition with hybrid ctc/rnn-t fast conformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.288628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.690215Z digest=sha256:8781657f9b8120a80d6ca76c073d67be6fd7feffcbeaa07a730c9226a66ba1c6

Observation f2c10cca-c52b-4b20-8197-d95c478c2fa3 · outbound

This paper cites Lrs3-ted: a large-scale dataset for visual speech recognition, 2018.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lrs3-ted: a large-scale dataset for visual speech recognition, 2018

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.695123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.695123Z digest=sha256:c88a8f92374bacd49ae9e7e7b3c8ebff6163a99976eb75fd33e393d163fd47de

Observation c6b12aa4-1e3a-49f9-98c3-f0fe336b05b8 · outbound

This paper cites MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.699358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.699358Z digest=sha256:62be3fffbac572dcdd0976832df9c698508ce7d270209626f8f642c2c0fc9e8f

Observation 9f401350-0f5e-4d95-9617-f9bb28436777 · outbound

This paper cites Braven: Improving self-supervised pre-training for visual and auditory speech recognition, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Braven: Improving self-supervised pre-training for visual and auditory speech recognition, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.263968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.703772Z digest=sha256:ec86646a692ae3084d305da88cf10df2a39f9718610a98692c61f99de546a65c

Observation 7e0f453d-368b-4502-bfa2-bcdfc91b8238 · outbound

This paper cites Large language models are strong audio-visual speech recognition learners, 2025.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Large language models are strong audio-visual speech recognition learners, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.708017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.708017Z digest=sha256:7cf849e0932ddfa2d48ef1e840d4a98790bfcab55322986d48a7e9301fe63385

Observation a4c61101-e1e3-4c65-bfe4-7ef2d5ad22b7 · outbound

This paper cites Visualvoice: Audio-visual speech separation with cross-modal consistency, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Visualvoice: Audio-visual speech separation with cross-modal consistency, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.238909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.712253Z digest=sha256:71110ea5861a1d69d7d2b1e9c3777515de2405a0b2757ac8ee6ca4904e05f6d8

Observation b9740d49-4e6c-4702-b84c-f0d4c86339d1 · outbound

This paper cites Ctcnet: A cnn- transformer cooperation network for face image super-resolution.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Ctcnet: A cnn- transformer cooperation network for face image super-resolution

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.224269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.716572Z digest=sha256:c9fb5584e67279f0122f2c9910c9fba0be61deeb262e6094ff7d3886d9521299

Observation 5127e6c6-1d23-4421-ad8e-df41f411260e · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Muse: Multi-modal target speaker extraction with visual cues, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.209791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.720990Z digest=sha256:61cbf8f83a4c4466e33dbaff5d3ea0a429dfbc198168f66f136be157aef5c66c

Observation 4de8b763-bce7-4c53-9f63-e2b381b24b73 · outbound

This paper cites Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.193757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.725497Z digest=sha256:e664684da357ed2a9f47d0e2c081178131290419869b375b3121b2ec069e5890

Observation 86797946-3f20-476d-9058-dc0a29b65927 · outbound

This paper cites Attention is all you need in speech separation, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Attention is all you need in speech separation, 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.179881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.729752Z digest=sha256:bcc75191e599e3fbdbae666201afde6f84795efb1318f6114b40f2ea09b244e9

Observation 2306f3dc-d4e5-4afd-b244-0d46e6daa9c1 · outbound

This paper cites Separate in the speech chain: Cross-modal conditional audio-visual target speech extraction, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Separate in the speech chain: Cross-modal conditional audio-visual target speech extraction, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.163493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.734164Z digest=sha256:1f168004685c07150a283bbf78f511a6bdd3fdcb21e5244955bd3f216e6c39e5

Observation 42cc8e8e-7598-4613-b881-5b0a41f0a44a · outbound

This paper cites A light weight model for active speaker detection, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization A light weight model for active speaker detection, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.147444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.738490Z digest=sha256:c7e040abb91ef6c67267ceb477b51bcc0e23fd33bc33b818415931ef251eb7ae

Observation efdd07a1-4504-4b6e-9538-2dfe98b7b255 · outbound

This paper cites End-to-end active speaker detection, 2022.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization End-to-end active speaker detection, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.132880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.743048Z digest=sha256:3f87280c42d2bec343710c3bde0d423ada4518318e58e4655774ae581d848aae

Observation f4bdcfad-cf29-4f35-bd6d-5257cdf90f1f · outbound

This paper cites Loconet: Long-short context network for active speaker detection, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Loconet: Long-short context network for active speaker detection, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.118510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.747136Z digest=sha256:abfbb6ed30e001a27a3d8d608f06d2e80f8fe8f2212740f3336f50128ad60b68

Observation 961a5d6d-51b1-41b1-af05-564ebfbbf10b · outbound

This paper cites Lr-asd: Lightweight and robust network for active speaker detection.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lr-asd: Lightweight and robust network for active speaker detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.104097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.751398Z digest=sha256:0fee0fe257c03157d791c1d21556c8d03d9c649223d26d57421bf28a97bf0cef

Observation ea3bdbd8-7c68-4200-a748-65a4d799d393 · outbound

This paper cites Tdn: Temporal difference networks for efficient action recognition, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Tdn: Temporal difference networks for efficient action recognition, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.089734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.755649Z digest=sha256:44a78601be630636696f089c7c44c38e0a979867d16878559f3ba91cfc4bc603

Observation f3be7b5b-adc9-468b-9668-f40f52bc606d · outbound

This paper cites Dauphin, Angela Fan, Michael Auli, and David Grangier.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Dauphin, Angela Fan, Michael Auli, and David Grangier

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.759904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.759904Z digest=sha256:565644d3090e7ce19d5ba778d062a9845b5a48c746f9138a3737b7831b84e85d

Observation 23630693-a3ce-4533-b4ae-8efd72aa2bd1 · outbound

This paper cites Musan: A music, speech, and noise corpus, 2015.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Musan: A music, speech, and noise corpus, 2015

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.764101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.764101Z digest=sha256:35de5ea5dc19ad5f15c1cc96272cb522b010b971971f4ca4305bc87e06e45f60

Observation 04d071a2-fb16-4b7b-bb6d-b6a310bdfb8d · outbound

This paper cites Audio- visual speech recognition with a hybrid ctc/attention architecture, 2018.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio- visual speech recognition with a hybrid ctc/attention architecture, 2018

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.768340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.768340Z digest=sha256:2d61b815c710a51f56f4c48c7516d75f1ab2d003f87ae0a7578cca1f98a79574

Observation c07be99f-6f8c-4ca5-b7a2-81137f33c9d7 · outbound

This paper cites Es3: Evolving self-supervised learning of robust audio-visual speech representations.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Es3: Evolving self-supervised learning of robust audio-visual speech representations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.045641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.772581Z digest=sha256:ff878dcd2f83d490f5acd886988492950cadb44ef587570fe6cd546ca288f7c2

Observation 73c5d417-7494-4daf-940f-cf1880addee0 · outbound

This paper cites Syncvsr: Data-efficient visual speech recognition with end-to-end crossmodal audio token synchronization, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Syncvsr: Data-efficient visual speech recognition with end-to-end crossmodal audio token synchronization, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.031534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.777319Z digest=sha256:abecdecc377c8730998a6bdd6e3b5b7c4c22f129765768db086f8226d9cacae4

Observation 6d3df58b-f8ac-4d6b-a411-aca47a226325 · outbound

This paper cites Sub-word level lip reading with visual attention, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Sub-word level lip reading with visual attention, 2021

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.016589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.781608Z digest=sha256:99af0d418635038bfa0a59deb36446761b90ecb99241a1c8ed927fdea728190b

Observation d5b44847-8856-4c5c-98f2-8588f060171a · outbound

This paper cites Deep audio-visual speech recognition.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Deep audio-visual speech recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.000505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.787191Z digest=sha256:a38b60c0c82987b2bb12e60eab14635236a79c6257f9526e059ba26238646a9e

Observation 88998e02-a846-4a4b-80e2-f2dce41aa850 · outbound

This paper cites Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition, 2022.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.986821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.792140Z digest=sha256:e428656ffe1806756b151a9b9759854d7f57f2e9e82664cf91296804fe55cbfe

Observation d474d59e-b104-4927-8778-875557e3b091 · outbound

This paper cites Audio-visual efficient conformer for robust speech recognition, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio-visual efficient conformer for robust speech recognition, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.972183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.796585Z digest=sha256:ccc4797900f048ff20f5c03df04cb3902a1393e0fc95b29c57d4cdca39c2b836

Observation d8750052-a914-4197-980a-68f7e040c139 · outbound

This paper cites Audio-visual speech enhancement and separation by utilizing multi-modal self-supervised embeddings, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio-visual speech enhancement and separation by utilizing multi-modal self-supervised embeddings, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.957484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.801054Z digest=sha256:a040bbb6f2e7a374fbef0765768927f10313dcc40bb88b411c0351a55145cb9a

Observation 1412804f-e3a4-46f8-be06-6b245d5192ae · outbound

This paper cites Time domain audio visual speech separation, 2019.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Time domain audio visual speech separation, 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.942609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:02:14.805352Z digest=sha256:0fce74e0f09e2b4b42227524375af70ddca8146f89864c52c318c17461fb2231

Pith citing papers

No inbound Pith citation observations are available.