Pith. sign in

Paper Citation Record · LEDGER

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.03964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03964 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:38:54.409813Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 112c890a-9b33-4f67-b96c-f88dd6add5c8 · outbound

This paper cites ACM Transactions on Graphics , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE ACM Transactions on Graphics , volume =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.211546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.211546Z digest=sha256:9e627180bc0897bd3d1f1814c382f52b1b671ea1e115e44c96ab0e802d443b7a

Observation 1b927815-17bc-4886-93d9-8328dba9f8f2 · outbound

This paper cites 2025 , publisher=.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2025 , publisher=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.216918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.216918Z digest=sha256:5d0a2d6f3af5fe79a9c9fecad6c566662bea6b92185e9719d3b6de62fe5148c8

Observation 7e894187-9100-46a6-8547-10060c2cb37e · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.220345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.220345Z digest=sha256:a8b8612f3b06915b9ebae253045c4a6c57fc58eb71f8b25268230344d920adb6

Observation a800df13-33f0-4c3c-a4f2-850072516a9c · outbound

This paper cites Interspeech , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Interspeech , pages =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.223368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.223368Z digest=sha256:897cb0c508ea2d22574873c95a67401fc1c26cfe01368bf5c0e26619deae6705

Observation 61dd3564-683a-450b-845d-d1287cf94a61 · outbound

This paper cites 2022 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2022 , doi =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.226354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.226354Z digest=sha256:3232025904b2acb2f2ba0f8a43696c01b78c581b9d3827c9f57cb9ec23b82c2d

Observation 60a2f955-fd1c-414c-84b4-e3f86d41a8db · outbound

This paper cites 2019 , publisher=.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2019 , publisher=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.230320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.230320Z digest=sha256:309ee383db8b62213dd99467ee02c654207d39554e47c133aecbe9930a343f46

Observation b1da064b-7657-4104-8e98-88b95502d472 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.234316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.234316Z digest=sha256:5f52d5fd2d33d5981be7236b3adb3a7c65eed3f1b082c75407d2b0307a8ae8b3

Observation 79a3f8dc-cd62-4314-9f00-1e1d96af9493 · outbound

This paper cites Interspeech , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Interspeech , pages =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.238054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.238054Z digest=sha256:3341253b5132844602da50cc2cd1b7c85dd699ff31653af22e780ffb8e9dae65

Observation 20b3a242-fc20-4c28-9ed8-f26fd7675079 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.241584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.241584Z digest=sha256:412900d6479135bd6b728c3631b2f2a3c047b57685c12c2dbcd4e227f32318aa

Observation 2194eccc-70b0-434a-b99e-f97e018babd5 · outbound

This paper cites 2018 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2018 , doi =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.245641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.245641Z digest=sha256:e01a7c22e9c196963dfd9673213f178405321ed2ce1d987de583629a7075d494

Observation d1d82b5c-84ad-4417-9f7b-419d06fa569a · outbound

This paper cites , booktitle =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE , booktitle =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.249341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.249341Z digest=sha256:75d8cb4ca78bef6e4520f658e93aef822b828cfce84d04d5e95291cb41dce888

Observation 539382fb-ce8d-46d8-a2ae-417fbefb4304 · outbound

This paper cites 2017 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2017 , doi =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.252885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.252885Z digest=sha256:ff1a682508db68a81b1229e4c0a7d22b6a709cf7047d8e9f1b63db966b739753

Observation 7a3a86a3-2fcf-41b5-9193-2af8c5f0e682 · outbound

This paper cites 2023 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2023 , doi =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.256292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.256292Z digest=sha256:110b260f0ee62668fd422fde435e9e73210055af16447a376c1bdc0a3a818b0d

Observation ad1c46d9-1b70-4424-8411-d6c996dd1d43 · outbound

This paper cites Scenario-Aware Audio-Visual.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Scenario-Aware Audio-Visual

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.259841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.259841Z digest=sha256:9d71ecf559e9cf86e00b197dfd3cd855e457527d4aeb8be7fbe3db1710ffeb85

Observation 1d007458-36a0-48fd-9d24-6c4607fd5b62 · outbound

This paper cites , title =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE , title =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.263337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.263337Z digest=sha256:9f45e25e65d60449e42a3097b9930cc261cf511d849e929e540b211dd14885a7

Observation e88cb523-481f-4f3f-be00-5b88767ba1c1 · outbound

This paper cites 2018 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2018 , doi =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.266927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.266927Z digest=sha256:be30d87d641ffdc28fdda07c1db0852743af58bbc8e3e6221a6f82aa5548782b

Observation 3f31e0d4-41af-40d5-93f5-3a3e0df90af0 · outbound

This paper cites 2020 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2020 , doi =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.270343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.270343Z digest=sha256:a01b4cd38e3aec1af248d6bd03f68c7937ed231d4601012758b66c234115e15f

Observation ef4f2bf9-77aa-4a9e-8de0-12b4b396d984 · outbound

This paper cites 2020 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2020 , doi =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.274039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.274039Z digest=sha256:b2607da07047a1988350c1c5b7f944f7a842b4f58d5dc0b88aa12f226b3e7c52

Observation 4266f7a6-6c7c-4505-9e9b-285d22f33640 · outbound

This paper cites IEEE International Conference on Computer Vision , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE IEEE International Conference on Computer Vision , pages =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.277716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.277716Z digest=sha256:a065d00b015103f017ab09f182fd310f032c2fabaa94dc39952c06b7bda75cba

Observation 6f8e3df5-751c-4209-8af3-da698735f802 · outbound

This paper cites 2015 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2015 , doi =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.281291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.281291Z digest=sha256:f58650b7fc2c6fb2252e2125316f0804b06f64d71000a89f9471f0f6559adaf1

Observation 807b4805-0e65-461e-aa90-f9a65874a224 · outbound

This paper cites and Zisserman, Andrew , booktitle =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE and Zisserman, Andrew , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.285144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.285144Z digest=sha256:e300c3651de31dfb430a51e0a59adbab983f80c22922e4c672d615e786625c9a

Observation ab9bbf30-5cc8-4f92-b1bd-abca76a9406e · outbound

This paper cites IEEE Transactions on Information Theory , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE IEEE Transactions on Information Theory , volume =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.289014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.289014Z digest=sha256:4911bd1a5020c71bba16d09e16b7c185ca264c3870702944e03787a3e2a435c0

Observation 5984d529-99f9-419b-a51a-12911f513f01 · outbound

This paper cites ACM Multimedia , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE ACM Multimedia , pages =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.292831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.292831Z digest=sha256:54ef83e76afaa65f870a02c6f63c5707e7c7f33da6ef27b9bd4230d66829bc32

Observation 28122e67-d504-4689-9c89-c85ca1dddeed · outbound

This paper cites The Annals of Statistics , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE The Annals of Statistics , volume =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.296240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.296240Z digest=sha256:40b9d6df56cbd5b2d90c291d7d9ddb6893e1acbc578dfa5e4b67beeaee5e6b7b

Observation a02291c4-9efc-460e-99ac-86008ad23aab · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.308651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.308651Z digest=sha256:649af94e95abe1f3598ecf5eabd23eeece600b5efeeb57cf9fa830607f13cadf

Observation 03294fde-dc69-438b-a96b-649434e648de · outbound

This paper cites 2026 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2026 , doi =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.311963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.311963Z digest=sha256:3fcf0b445c88f762bbd1a28e86ca9a599e9a2956db0e59e7c781de283dbcd47c

Observation 7b468ec1-be23-4dfc-8e6b-404de216cc77 · outbound

This paper cites 2025 , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2025 , pages =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.315406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.315406Z digest=sha256:310f4b3b08150733e31bd8e50018245996f58f563d231c0ec20b9b0a73d72a59

Observation b333e601-2da3-4327-8a19-d5cfc83df7de · outbound

This paper cites ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.318691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.318691Z digest=sha256:bbaf93d439ce9cf92e6e45ede129145ef617ab881334a2d1238e020ab7b7b4e3

Observation 86254853-121b-4cfb-a3a2-9fd481c3351c · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE LRS3-TED: a large-scale dataset for visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.321955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.321955Z digest=sha256:2fd7499f0013983947ea17ec94d354feb921f52d9c00da6ef170e7847060ae9d

Observation 474d28ca-f1d0-43b0-993b-49d471749e2c · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.325800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.325800Z digest=sha256:e8f087ed7fff0c63d18f41b22d8ce3f093b7964954c401bf4927534c2a833c37

Observation 0a328093-3103-4743-9785-21100acc16f8 · outbound

This paper cites Parkhi, and Andrew Zisserman.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Parkhi, and Andrew Zisserman

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.329420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.329420Z digest=sha256:a68b01956dab35d9a5d70cad1c6b218ee285cf814f17328b71a7115077a78790

Observation 65c05404-29c9-4730-bba4-30d59e2db9d7 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.333168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.333168Z digest=sha256:a7a907adfcc1a01795c1626b5b9c5bf812de393a6b3f9bff0e10f7fb02aa7066

Observation 8f61b9d5-5d86-4ab4-afb3-33107d51200a · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.336600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.336600Z digest=sha256:13872ec8ddaa5dcbe0b15586d8aa9eecf8d7c348d492091ae99f24218e528bb5

Observation 313acf9d-0870-4420-abe3-4b99a0d74da6 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.339648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.339648Z digest=sha256:9a38c04efd84b307a4b3846e1d3026800a1cc02b6b71b3a4bc4ccca097cac21d

Observation ea59e885-92f5-4560-a346-74d4ce5299b2 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.342520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.342520Z digest=sha256:6627c48a9f4a6b73357a4de28e00233a00c62e55dc2278a03fee0431b1c0fd6e

Observation bb2f8a39-2f60-4774-bd1c-231f8f3b3861 · outbound

This paper cites Freeman, and Michael Rubinstein.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Freeman, and Michael Rubinstein

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.345337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.345337Z digest=sha256:1600beeeba64b3fe62ddd01226d1624c53649089df5e4277e8ec0a772aa6bc0a

Observation 96e1c602-b0f3-4613-b820-20faf5b74c85 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 39

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T00:38:55.208961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T00:38:54.348174Z digest=sha256:4331499ae146c4f9018e5c8b694d4036323589c5dca3acb64be5b6765e1609b1

Observation 46ca842a-cc4c-49d1-9e69-0373a9275f21 · outbound

This paper cites Levenshtein.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Levenshtein

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.351032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.351032Z digest=sha256:bb865469d9d6b57305860c53a7a3faf7052ffc58301981ef7b431b568b204581

Observation f662acfe-81d3-45c6-b467-6f7ec3a4da23 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.353708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.353708Z digest=sha256:ad422a2953a4a9a1b2de71d231188b6bb65b11b3e2a0f08e20f992880695ffc3

Observation 044b5cc1-d2ee-45b1-9a6b-03240fad9dde · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.356936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.356936Z digest=sha256:557f30089736685f3c00170674b3f4b6cd7a6581a043c235029432d26b8eee3c

Observation 5ed23fc1-67b8-474b-ac58-dba8ecfaf0ca · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.359955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.359955Z digest=sha256:93c71d4460f759ad84f924c0d5fdec764f25430ec570972452e3c184fc0f82cd

Observation 042ccf02-0a6f-4e25-b1fe-615dfb96d307 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.363402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.363402Z digest=sha256:cd3e22afd39e043616bf96e1258a7079ab6a69ccc13d21e5203ab2b43f51e24c

Observation 61a2b157-5986-4c77-9182-5c40862a7f61 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.366914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.366914Z digest=sha256:fcf501211093deace0bbb8ed69f3ad6b3955e8ad8618f3b284e6afc03a20100e

Observation 92025b76-f9ab-4d82-832e-057e8e627e9b · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T00:38:55.094647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T00:38:54.370357Z digest=sha256:8b27a267a71f04e133b0c88e82713dff0a0e8c4c17677b977e4accf18806a0a4

Observation 94bb6707-9a05-4651-94e0-7ac82a880576 · outbound

This paper cites Germain, Sameer Khurana, Chiori Hori, and Jonathan Le Roux.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Germain, Sameer Khurana, Chiori Hori, and Jonathan Le Roux

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.373896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.373896Z digest=sha256:833e56e9317fa7db023f20b1cef0e6169cd2d3e4c96ae9b639bb506355c9bd5b

Observation 124d17c8-d2f7-4916-aa70-48a6b57d695c · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.377234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.377234Z digest=sha256:1e1033123ca511ac59de77e8a9a4edcb35ef59920be0d14b325aed936b6805a1

Observation 0308c74e-745e-4bdb-80a9-fc395a666122 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.380506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.380506Z digest=sha256:ef9b1bdcca3d26e21a4b55d82f0272a4455d6997eaddc779b480342a87df8077

Observation 137980d6-499a-466b-8246-37875d31c9a0 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.383895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.383895Z digest=sha256:9b2fc366a6c7e4770a51fca4a98f2a94057b61fcb24b904c87b83f3084f4e801

Observation f48c3caa-2414-4616-9134-fb48b647a39f · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.387173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.387173Z digest=sha256:94f2db506b7115ed878470b031b6b3faef83e1027e9b3ac7f02deafbd16e0880

Observation 5862b241-aec2-4ec6-929c-eb3016af799b · outbound

This paper cites Qwen3-ASR Technical Report.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Qwen3-ASR Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.390520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.390520Z digest=sha256:68f6dbaa1e57bb841174456cf1620fd4a793b608d68d05fde7dae8cd63651255

Observation 90aa7cfb-9ff0-4588-8bc7-9aca6531904e · outbound

This paper cites M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.394031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.394031Z digest=sha256:f1b89151e68513785f894ff7dfbeed9abf665e055f731345f3874a5fd90f3180

Observation 2524002a-08e6-4e81-966b-538796c011ba · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.397062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.397062Z digest=sha256:b4bd57f364398ade8753645cdf78c793800a0e52cf7e8ceb6af395a4836309e6

Observation f153512a-cb7b-4cd9-a6db-55261bd5f2c3 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.400282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.400282Z digest=sha256:b9d5131fbdb3bc0a1fcfd4970a22b9e6e578300845d387871526fcaecde75253

Observation b0693668-e234-4b58-b3d5-4e5d405ea64d · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 56

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T00:38:54.704054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T00:38:54.403395Z digest=sha256:2cef355603221f262f9b653115ee86df6a40a428a585db9a279ca37d95e35b8d

Observation 69a62879-b9c3-403c-8956-f050ccbad675 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.406547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.406547Z digest=sha256:38437e1b9ea4cbc45fddaf4c12c8c110577dd5608ed8ab232b7940915b663005

Observation 3fa79b04-247c-4abe-bdf6-87be5cef118c · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.409813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.409813Z digest=sha256:501ac0010988d7f53ae806d52bb85a5d521f27507e81c135dff7566b7f4bdad2

Pith citing papers

No inbound Pith citation observations are available.