Pith. sign in

Paper Citation Record · LEDGER

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

As of 9 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.03964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03964 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:38:54.409813Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 112c890a-9b33-4f67-b96c-f88dd6add5c8 · outbound

This paper cites ACM Transactions on Graphics , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE ACM Transactions on Graphics , volume =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.211546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.211546Z digest=sha256:4c54ddbfb863163ef4e7299cd3bebbb383b695307ab6eb70fc0801c1eb8455dd

Observation 1b927815-17bc-4886-93d9-8328dba9f8f2 · outbound

This paper cites 2025 , publisher=.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2025 , publisher=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.216918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.216918Z digest=sha256:e4f5251cebc22a0b60ffb685829d2487d3b24f221c5ca9b83a75a760b0771c1d

Observation 7e894187-9100-46a6-8547-10060c2cb37e · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.220345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.220345Z digest=sha256:2ae030ed9b0ea7a920de7a96801bd892379b147b043fe3e9f8d5061081468e26

Observation a800df13-33f0-4c3c-a4f2-850072516a9c · outbound

This paper cites Interspeech , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Interspeech , pages =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.223368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.223368Z digest=sha256:fe499c000515f8ee8552f142979663850e53a1a01c602fdf5e1097d02c0c6c5f

Observation 61dd3564-683a-450b-845d-d1287cf94a61 · outbound

This paper cites 2022 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2022 , doi =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.226354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.226354Z digest=sha256:55d2deeb975b6972f480411e682d76e831ecf6c9fcc1492be71953b4b118be59

Observation 60a2f955-fd1c-414c-84b4-e3f86d41a8db · outbound

This paper cites 2019 , publisher=.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2019 , publisher=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.230320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.230320Z digest=sha256:120d6f171f54c0075d1e65e78dd054b3f67adfd8d44601416beaadd0776a3f1e

Observation b1da064b-7657-4104-8e98-88b95502d472 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.234316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.234316Z digest=sha256:9d0637d840d88834189d68bcdc005a8fb7461ba0e5eed473db99c7feff8993d8

Observation 79a3f8dc-cd62-4314-9f00-1e1d96af9493 · outbound

This paper cites Interspeech , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Interspeech , pages =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.238054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.238054Z digest=sha256:b3fae5dda35c7a8c805432f4b2c0a5b9c5ffd7f3d457e8ffa33ae2f118e3319f

Observation 20b3a242-fc20-4c28-9ed8-f26fd7675079 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.241584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.241584Z digest=sha256:6d0c13fe529f401832e2fdbfda638f06ad814b94b662a81cecc8d288dca17d06

Observation 2194eccc-70b0-434a-b99e-f97e018babd5 · outbound

This paper cites 2018 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2018 , doi =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.245641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.245641Z digest=sha256:4bb252fac257d02dbb86ad8d2758106cd62236cb83267e1fb0c66f3b1840ceff

Observation d1d82b5c-84ad-4417-9f7b-419d06fa569a · outbound

This paper cites , booktitle =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE , booktitle =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.249341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.249341Z digest=sha256:dad6bf314e828b56df2c73da73b9fa90968aa8861f17a96042487767329d9c2a

Observation 539382fb-ce8d-46d8-a2ae-417fbefb4304 · outbound

This paper cites 2017 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2017 , doi =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.252885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.252885Z digest=sha256:39baae7cd12d8806b9ef62d8f1161e59bdf7dbdcc6cbe4f698ad4cf41d08fdf4

Observation 7a3a86a3-2fcf-41b5-9193-2af8c5f0e682 · outbound

This paper cites 2023 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2023 , doi =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.256292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.256292Z digest=sha256:6542263dbf91e30b4a5865799cb8675ec7ff4938f9168e001d0fd9c72072c93b

Observation ad1c46d9-1b70-4424-8411-d6c996dd1d43 · outbound

This paper cites Scenario-Aware Audio-Visual.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Scenario-Aware Audio-Visual

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.259841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.259841Z digest=sha256:893da6e81af5812e00981f089567fb118424d73a7219c2b839c04bcdb1349ca3

Observation 1d007458-36a0-48fd-9d24-6c4607fd5b62 · outbound

This paper cites , title =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE , title =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.263337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.263337Z digest=sha256:212dd2ed69dadad848d709859cf1514b39629c23e726003dd9a2d33380d6d563

Observation e88cb523-481f-4f3f-be00-5b88767ba1c1 · outbound

This paper cites 2018 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2018 , doi =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.266927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.266927Z digest=sha256:fcba25d1855432ad2239f1177c9ca30c98d73fa66638af4d762a825cd72a2bb7

Observation 3f31e0d4-41af-40d5-93f5-3a3e0df90af0 · outbound

This paper cites 2020 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2020 , doi =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.270343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.270343Z digest=sha256:eef1fa824e6342de769ccb9b55144da2e87f2d28fc20c1d6ddb00078a952e05b

Observation ef4f2bf9-77aa-4a9e-8de0-12b4b396d984 · outbound

This paper cites 2020 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2020 , doi =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.274039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.274039Z digest=sha256:f6ccd5ac284a6729258bf466d7ff8d1cb41c9a2cd0efe083d0ec060c214cce67

Observation 4266f7a6-6c7c-4505-9e9b-285d22f33640 · outbound

This paper cites IEEE International Conference on Computer Vision , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE IEEE International Conference on Computer Vision , pages =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.277716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.277716Z digest=sha256:f4273cb609091937946fe07ed97aabfbd182646b8eb4567ba00028d6aef3da14

Observation 6f8e3df5-751c-4209-8af3-da698735f802 · outbound

This paper cites 2015 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2015 , doi =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.281291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.281291Z digest=sha256:bc60cd960df3f21e1633fd265fc57daa0bcc2cb0acce9cd2b50a0641ecda9a61

Observation 807b4805-0e65-461e-aa90-f9a65874a224 · outbound

This paper cites and Zisserman, Andrew , booktitle =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE and Zisserman, Andrew , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.285144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.285144Z digest=sha256:e3705261312dbebc15976b9e3d46ebbcfab168c902a6772dff32898f365eb682

Observation ab9bbf30-5cc8-4f92-b1bd-abca76a9406e · outbound

This paper cites IEEE Transactions on Information Theory , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE IEEE Transactions on Information Theory , volume =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.289014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.289014Z digest=sha256:0788113aca8913e44b6efbea0d4090cb35b9e7f802bccdad3500e422f5ddb94e

Observation 5984d529-99f9-419b-a51a-12911f513f01 · outbound

This paper cites ACM Multimedia , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE ACM Multimedia , pages =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.292831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.292831Z digest=sha256:8f964d7e4749de0251ff80a582d27f0944e13b284be5339ab3b4fe82889a164b

Observation 28122e67-d504-4689-9c89-c85ca1dddeed · outbound

This paper cites The Annals of Statistics , volume =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE The Annals of Statistics , volume =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.296240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.296240Z digest=sha256:bc5662f91cd3265bfe2f63ceae2edb2ee40ee979b1236640a3c219a69735b8a9

Observation a02291c4-9efc-460e-99ac-86008ad23aab · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.308651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.308651Z digest=sha256:39defaa166de422d995d56b17982adf8fb4c099e054a800314037b559ed1d72e

Observation 03294fde-dc69-438b-a96b-649434e648de · outbound

This paper cites 2026 , doi =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2026 , doi =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.311963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.311963Z digest=sha256:20017c56f1d14c2369a5f4e3f7519331521cd13b429e6f3457daeca97703ade3

Observation 7b468ec1-be23-4dfc-8e6b-404de216cc77 · outbound

This paper cites 2025 , pages =.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE 2025 , pages =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.315406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.315406Z digest=sha256:3545f6c331e8ecb4ed2c6fd821ade651147431f87d64dca6526bc5619c1a2a38

Observation b333e601-2da3-4327-8a19-d5cfc83df7de · outbound

This paper cites ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.318691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.318691Z digest=sha256:d753a1ef239e7cc2e7d16f48dc7bf391217ce8b0fd848d8b80da2916ed94d25c

Observation 86254853-121b-4cfb-a3a2-9fd481c3351c · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE LRS3-TED: a large-scale dataset for visual speech recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.321955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.321955Z digest=sha256:a39a973aa89d733ad1b0ca7a3ba580268078eafaac9847a405b6c15cb53f381b

Observation 474d28ca-f1d0-43b0-993b-49d471749e2c · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.325800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.325800Z digest=sha256:69a292ab02ccdaa66d66b627077b118333840c0044e31fb6aa494b963a81a184

Observation 0a328093-3103-4743-9785-21100acc16f8 · outbound

This paper cites Parkhi, and Andrew Zisserman.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Parkhi, and Andrew Zisserman

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.329420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.329420Z digest=sha256:6ff93b0667c76a8cb1b4837b7d39791892838d499721e0d842d5542a2b7af8d4

Observation 65c05404-29c9-4730-bba4-30d59e2db9d7 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.333168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.333168Z digest=sha256:2ed4b0e2f0bfc09087b72a3f549ebcaa60d839af0558bd54473e780002dd7ebb

Observation 8f61b9d5-5d86-4ab4-afb3-33107d51200a · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.336600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.336600Z digest=sha256:3580554a05a5c04cc6659627644c35def94dcc7ccbf002bd1cd98a0f8d4148ff

Observation 313acf9d-0870-4420-abe3-4b99a0d74da6 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.339648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.339648Z digest=sha256:3f3c42a43441b974ca1bcb57399465a7eaace49f2307fdbe7e06a25cfe8244b7

Observation ea59e885-92f5-4560-a346-74d4ce5299b2 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.342520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.342520Z digest=sha256:02dd849575fdc57f5fcc739b94194d3f1e4dad8b9db08ff66a209a69094570a0

Observation bb2f8a39-2f60-4774-bd1c-231f8f3b3861 · outbound

This paper cites Freeman, and Michael Rubinstein.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Freeman, and Michael Rubinstein

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.345337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.345337Z digest=sha256:a908fdf81629fe8449a902906d5ed783fff1a77a9aca3da2fdf8103c49afe0e7

Observation 96e1c602-b0f3-4613-b820-20faf5b74c85 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 39

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T00:38:55.208961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:38:54.348174Z digest=sha256:d9d877f27937799d2039c54d2ef0c58f0a8e4be969d0295851cc81b2d263e4af

Observation 46ca842a-cc4c-49d1-9e69-0373a9275f21 · outbound

This paper cites Levenshtein.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Levenshtein

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.351032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.351032Z digest=sha256:a20b7add4bfcc928d052d0921779654d57885fca67542760728fb4f9b16546a8

Observation f662acfe-81d3-45c6-b467-6f7ec3a4da23 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.353708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.353708Z digest=sha256:2bea93d1146d236e809156074d1746923d4a08ba299355b64c957572c17ff8a1

Observation 044b5cc1-d2ee-45b1-9a6b-03240fad9dde · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.356936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.356936Z digest=sha256:d4a2f91f969d9597df2a3aa4650def64ec34b0f97997be66b437802d3689a48b

Observation 5ed23fc1-67b8-474b-ac58-dba8ecfaf0ca · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.359955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.359955Z digest=sha256:ddfbc00cb4bdef11796866c0a82f4a213434ef0ea3933b6775202fd51496ad96

Observation 042ccf02-0a6f-4e25-b1fe-615dfb96d307 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.363402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.363402Z digest=sha256:b995c66586ef4da239fb8229bb6b5e533ff9a301dc4d1f5ed5b1baee80ed6c46

Observation 61a2b157-5986-4c77-9182-5c40862a7f61 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.366914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.366914Z digest=sha256:9f7ec8bf7bea2e7d613caba22736d919580b679543b7c8b5c620961c062290f9

Observation 92025b76-f9ab-4d82-832e-057e8e627e9b · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T00:38:55.094647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:38:54.370357Z digest=sha256:47afbceff9fefd16e61b4808dd3dfad36caced9bd0f4f28b1feffb3a81790d1f

Observation 94bb6707-9a05-4651-94e0-7ac82a880576 · outbound

This paper cites Germain, Sameer Khurana, Chiori Hori, and Jonathan Le Roux.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Germain, Sameer Khurana, Chiori Hori, and Jonathan Le Roux

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.373896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.373896Z digest=sha256:61949e792ffe78b3b8b10e1fd2eb51f1d01cecd73ffba1b9199981a862b3c971

Observation 124d17c8-d2f7-4916-aa70-48a6b57d695c · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.377234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.377234Z digest=sha256:6e4a9be1c4e0113b378c9afc37e2109395ca546b55032629b7e193c38bdb7ce5

Observation 0308c74e-745e-4bdb-80a9-fc395a666122 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.380506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.380506Z digest=sha256:421897de136795fe047bae49917322227d7caf25de236cf6be5de93500189771

Observation 137980d6-499a-466b-8246-37875d31c9a0 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.383895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.383895Z digest=sha256:4bf42f4ad6d4ba882d8c5a9c2d50f7ee555bc0dd0066eedad2d3d9b6905b4946

Observation f48c3caa-2414-4616-9134-fb48b647a39f · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.387173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.387173Z digest=sha256:fa6ec258f2943d783e7e1b0691280d4cf25b35f17cbe324ae7a13809d295c229

Observation 5862b241-aec2-4ec6-929c-eb3016af799b · outbound

This paper cites Qwen3-ASR Technical Report.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Qwen3-ASR Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.390520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.390520Z digest=sha256:ef572637eff742fb62eaea8ee15df930aa658d1547bc5708e9a7e2485689f03c

Observation 90aa7cfb-9ff0-4588-8bc7-9aca6531904e · outbound

This paper cites M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.394031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.394031Z digest=sha256:bcca1a501722e14c0d9825dc5e1809b699fd0dbff67a70b4815a84a81efef861

Observation 2524002a-08e6-4e81-966b-538796c011ba · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.397062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.397062Z digest=sha256:2627a62c2583ba057096e97831b1ced568e834d3d2cf6545d1424f3ca4117303

Observation f153512a-cb7b-4cd9-a6db-55261bd5f2c3 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.400282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.400282Z digest=sha256:b82657f91d41d1aa3aac2910a2144663b2eab77ba48b0c5636c6a38044516c17

Observation b0693668-e234-4b58-b3d5-4e5d405ea64d · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 56

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T00:38:54.704054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:38:54.403395Z digest=sha256:7a4b8e711531aa8f14dd1406f8bf41d75a6eaba93b1731b55009c7570922040f

Observation 69a62879-b9c3-403c-8956-f050ccbad675 · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.406547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.406547Z digest=sha256:ecde6ee7559006943a2745786565166344ff701a838f796af13e9c5076bbe4ff

Observation 3fa79b04-247c-4abe-bdf6-87be5cef118c · outbound

This paper cites an unresolved cited work.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.409813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.409813Z digest=sha256:b130f43366228dc4cfcb08b2472cbda9ba4b8e65a967a92afa7c97f5719218c8

Pith citing papers

No inbound Pith citation observations are available.