Pith. sign in

Paper Citation Record · LEDGER

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.09792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09792 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.812385Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.622698Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:46:04.891505Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc2de33d-730c-49d9-8dc4-cc305f4e20d3 · outbound

This paper cites Most existing studies focus on improving the audio-visual fusion mechanisms [1, 2, 3, 4, 5] or addressing visual cue-impaired scenarios [6, 7].

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Most existing studies focus on improving the audio-visual fusion mechanisms [1, 2, 3, 4, 5] or addressing visual cue-impaired scenarios [6, 7]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.469713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.617264Z digest=sha256:c00ec6a439730c02d4dfc0bb37b73194f6e2bb0a02ea9c8bd96c76d6a01cb466

Observation bd51abf4-a74b-46fe-be2e-d77959f7cc4a · outbound

This paper cites Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:04.897605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.622698Z digest=sha256:15c7b3898c6788681ee7021a397de37b63012b5317198294995b28e8b3b41d3a

Observation bdec9dcb-8ec5-42ea-9e00-963626b2f9bd · outbound

This paper cites Dataset In this study, several experimental settings are considered: •Training Set:A two-speaker mixture training set is simu- lated following previous work [1, 6, 2, 3].

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Dataset In this study, several experimental settings are considered: •Training Set:A two-speaker mixture training set is simu- lated following previous work [1, 6, 2, 3]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.453563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.627656Z digest=sha256:6a7ab6a05d27534634cc3569dca238a3aef6af280e905bebc52fb51ed8260604

Observation 6dfff320-741e-4807-8bea-7aaa4718b772 · outbound

This paper cites an unresolved cited work.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:46:05.438722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.633169Z digest=sha256:8453ddc30aaf1cf3a8472cd8abf7ab0b7568ca5aed6d4b8397d7a97b0403c95e

Observation 715afddc-d228-491b-bd38-b22495bee4f9 · outbound

This paper cites Full occ.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Full occ

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:46:05.423908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.638070Z digest=sha256:b523cdbeb3ee9afec15e479a18740d144a1faf8109688f62a543707efe748308

Observation 554b79fa-87ed-4956-9e48-178c2cbd62e7 · outbound

This paper cites an unresolved cited work.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:46:05.408576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.643723Z digest=sha256:ed0664e2bd1a19d0679cbd424e06fda33507560f11be334f405941b56592e7b1

Observation f77eb3b3-5e62-42a7-a8a1-15d631fae36a · outbound

This paper cites 62401377, Shenzhen Sci- ence and Technology Program (Shenzhen Key Laboratory, Grant No.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction 62401377, Shenzhen Sci- ence and Technology Program (Shenzhen Key Laboratory, Grant No

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.393041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.648600Z digest=sha256:65fc8212e4f5d9c4721fc7d0a3cda28f2c0b0468188fdb36671090220973087c

Observation 277ee2e9-90dc-421c-8b38-bbdb74a5141e · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Muse: Multi-modal target speaker extraction with visual cues,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.377153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.653741Z digest=sha256:a9c2dc44bae1485d8322cb5986ab898635786b784ea73b13d3ed2728ccdcbfcf

Observation 5ecfee67-d124-4cc9-b4a4-b30340e60df6 · outbound

This paper cites Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.361888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.658898Z digest=sha256:8567bdac4a6f47e729f01ce894f858dd5e0d7018874e1c0e856931bd5e2a99ba

Observation dc1f0523-48b4-4580-b3cb-863bb54eea93 · outbound

This paper cites Avhumar: Audio- visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Avhumar: Audio- visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.347080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.663935Z digest=sha256:35b9572f9aab826149a866890eaf22cd0af9952c653a174d7d7d9f8bd078a420

Observation e9d1b61d-adb8-4864-84af-404c35a7c337 · outbound

This paper cites Target speech extraction with pre-trained av-hubert and mask-and-recover strat- egy,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Target speech extraction with pre-trained av-hubert and mask-and-recover strat- egy,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.331829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.669226Z digest=sha256:6426cc94795e5a5dfdd34379d1916520408fefaeac8190eadd677f15b029e342

Observation 9f2fce59-8df6-46fc-a264-defe611a2a0e · outbound

This paper cites c 2av-tse: Context and confidence-aware audio visual target speaker extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction c 2av-tse: Context and confidence-aware audio visual target speaker extraction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.317657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.674131Z digest=sha256:c9f9dcfc555439623a092c9af5ea7c00c49253f44e92bced2c00ba8053d7b059

Observation fb5b5b21-c426-4760-ba87-16c0ddf158b4 · outbound

This paper cites Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.303072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.679148Z digest=sha256:c759f27a13cc72c1bd4ca9342abb1267a1d224c19e8d51b731d4035cc453105b

Observation e0ad3710-a4cd-4595-8c2f-a4944678ed28 · outbound

This paper cites Restoring speaking lips from occlusion for audio-visual speech recognition,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Restoring speaking lips from occlusion for audio-visual speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.287472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.684022Z digest=sha256:33dac25126973d5f27f1d1e48959ac86af737d133301d55dfe5f719e9ddacbd0

Observation 1d9e7c30-6caf-405d-8f3f-df8ed062ab1f · outbound

This paper cites Semantic en- coding during language comprehension at single-cell resolution,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Semantic en- coding during language comprehension at single-cell resolution,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.271997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.689743Z digest=sha256:41a3576dea90dd9580ba027692bafa5f2712c98ebe569b9c7a4a419462dc1673

Observation 747ba67d-bf1f-43dd-bba5-3db1cc01ab66 · outbound

This paper cites Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.256584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.694126Z digest=sha256:12f8745749ed91ebfcbc6a9115de9d9c69b59faa7ac2bc150ab3751ba70987e1

Observation ce3d5104-6aa5-4de9-8092-a6289ee7cd5f · outbound

This paper cites Large language model can transcribe speech in multi-talker scenarios with versatile instructions,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Large language model can transcribe speech in multi-talker scenarios with versatile instructions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.240980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.698793Z digest=sha256:182fc23bae07e6c427c1e263beddbcff0849760b53cd5bbbb8bc5b06d458012f

Observation e133f23a-cfdd-4793-945f-bc8822d23e76 · outbound

This paper cites Target speech extraction with pre-trained self-supervised learning models,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Target speech extraction with pre-trained self-supervised learning models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.225837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.704344Z digest=sha256:d5cad1ae5d2206df84de763affd467a1bcf1497256d10e137ef6e67c97013f58

Observation e15dd3bf-0e42-49e5-aeb5-e76d5bcad2a6 · outbound

This paper cites Probing self-supervised learning models with target speech extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Probing self-supervised learning models with target speech extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.210381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.708724Z digest=sha256:b10013143db00a1a06c3b1f38c6276a52908f76675c2328a3c9a501eba5d3fbd

Observation 78c595a9-fa23-41e9-b7e4-e1ed7ad327c5 · outbound

This paper cites A large-scale evaluation of speech foundation models,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction A large-scale evaluation of speech foundation models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.194831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.713437Z digest=sha256:6658e7bc7c8b135a5f42d24aa4b96e68357f4f79dac578554337082e320693f7

Observation a571c4ce-40a4-4f71-b03d-16c5d536b767 · outbound

This paper cites Transferring knowledge from large foundation models to small downstream models,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Transferring knowledge from large foundation models to small downstream models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.179377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.718269Z digest=sha256:2718bc279e6dd94f44777f5b6a8957f4192362db92cb57a165790a774db73b81

Observation 62790262-403f-4144-b32a-f3b86abd677e · outbound

This paper cites Knowledge transfer from pre-trained language models to cif-based speech recognizers via hierarchical distillation,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Knowledge transfer from pre-trained language models to cif-based speech recognizers via hierarchical distillation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.163405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.723757Z digest=sha256:a2d21517879338e1f130d142af5df773e2789590bd9ae9a38b20dbf21c038e00

Observation b915031b-a18a-458b-90d7-4e2dd441ec03 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.146935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.728231Z digest=sha256:9b4aff0f1ca159a94f0ac5fdf7e39e6052efaa8331dc431f5152ec44cf295590

Observation 97115ffb-4df0-4dc0-9975-36526792fae7 · outbound

This paper cites LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.733824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.733824Z digest=sha256:6e0ae177347b8cf5a58e8cabd070473def29a1ec6beccb08d0d03b1f01f61076

Observation 2ddd5d4f-bc6b-4752-958a-3c196629d5fd · outbound

This paper cites ALMTokenizer: A Low- bitrate and Semantic-rich Audio Codec Tokenizer for Audio Lan- guage Modeling,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction ALMTokenizer: A Low- bitrate and Semantic-rich Audio Codec Tokenizer for Audio Lan- guage Modeling,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.131868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.738953Z digest=sha256:3f27e2e32359877ed07060c68e4ee43166cfccd45d8905e620d7fdaac4248387

Observation 4a37c920-5726-402b-9940-466c6bb4effb · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.116365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.744419Z digest=sha256:c755240b64261a8373dbf4fb15699a06f5ba75e404fc1819021c1175433387cf

Observation c2378fd2-e3ac-4f81-96a9-249ff3ea9aac · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Roberta: A robustly optimized bert pretraining approach,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.100660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.749038Z digest=sha256:da90affa25b07627726bb23c908771685b678fb4d2b826c540e1eb54b60bff42

Observation cccba63c-671b-4ce4-a72c-1fa36577a663 · outbound

This paper cites Separate in the speech chain: cross-modal conditional audio-visual target speech extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Separate in the speech chain: cross-modal conditional audio-visual target speech extraction,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.084425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.754727Z digest=sha256:f6afd9cbd42ec6ec22529ac36c7b956c8ece127d3a5813e23c6e23b8421a4715

Observation 9e63d60d-5eac-43ff-b8f5-113ff6fe942a · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction V oxceleb2: Deep speaker recognition,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.759101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.759101Z digest=sha256:6e4510586e7d1dfd78511e7f886179ebe57a6fd60051486f0a09480b8c6b1e4b

Observation cf384b2e-344f-4da7-a016-6670a52231d2 · outbound

This paper cites Watch or listen: Ro- bust audio-visual speech recognition with visual corruption mod- eling and reliability scoring,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Watch or listen: Ro- bust audio-visual speech recognition with visual corruption mod- eling and reliability scoring,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.058171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.763651Z digest=sha256:9957f674aa3c40c2c11357ffa3a5790be6041d272a88d73f9cd53f33db929a5f

Observation 8f3733b2-e3a8-4f34-9193-24faa3e99dcd · outbound

This paper cites Lrs3-ted: a large- scale dataset for visual speech recognition,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Lrs3-ted: a large- scale dataset for visual speech recognition,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.042463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.768090Z digest=sha256:e16fb2be230039617ad0fb925415cbff354dfc1a8de581464d090d03a132a94f

Observation df70e485-3872-4cd2-b00f-2584726a6d38 · outbound

This paper cites Sdr – half-baked or well done?.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Sdr – half-baked or well done?

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.027505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.773088Z digest=sha256:254836a5fd12e95d57d3782b93a265afcebabef760940ba6cdd5c972bcb0c092

Observation d35c370d-25fb-4aa2-92c0-6f0d5cbb1174 · outbound

This paper cites Single-sided Real-time PESQ Score Estimation.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Single-sided Real-time PESQ Score Estimation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:04.857164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.777432Z digest=sha256:a8a28c3f327422681b3af4b66cbbe77974ce31268cacc1df7d061a65a98615bb

Observation d0040916-1427-4b42-8027-a37154c8e2ae · outbound

This paper cites An al- gorithm for intelligibility prediction of time–frequency weighted noisy speech,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction An al- gorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.011753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.782208Z digest=sha256:74878e51b5268414ea44e79295977e73e73c2a13d575d3c21d371970dae94c6c

Observation 2feebede-1b54-43ca-a246-0f0eb293d75e · outbound

This paper cites SpeechBERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction SpeechBERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.995934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.787637Z digest=sha256:637454ae544109d128c5266539a5151f05b30ec5b7e3184f9d52568fa781d4f8

Observation e82f49b8-00ac-4725-ad93-a1c3d3503b1e · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction How should we extract discrete audio tokens from self-supervised models?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.979381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.792441Z digest=sha256:64ea56fe7fa3dba6248655764c6a47c7ab19bb223df57cfe155012ca076e8383

Observation 673af5ad-5b81-4606-a91d-8a5f8194060c · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.962900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.797897Z digest=sha256:87bc4ac42d6304d3911a13a85e31fdcec46750d5795b4f0ec0342c2c985d0664

Observation 5426972f-cbe5-4462-b88d-d586d4b961b1 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.946664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.802382Z digest=sha256:9d4b9103d5f296481ffeacc7124d2e0c488671f239cf5f853f2fd70d407dcac1

Observation c5af5730-58c1-48f1-870e-a77dab4ef6d3 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.929645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.806966Z digest=sha256:2fd12e05895dd07e32bb5e1225b074d92b7199500c774fc0d75d049c7807f42a

Observation 4ac4cd68-4834-42c4-8b31-c4753ab3a989 · outbound

This paper cites Intuitive multilingual audio- visual speech recognition with a single-trained model,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Intuitive multilingual audio- visual speech recognition with a single-trained model,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.913816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.812385Z digest=sha256:b74aa655e89f3dcb68dcea1a2beb8ec77f1a37772ab8adffda6bce242c87574a

Pith citing papers

Observation bd51abf4-a74b-46fe-be2e-d77959f7cc4a · inbound

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction cites this paper.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:04.897605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:46:04.622698Z digest=sha256:15c7b3898c6788681ee7021a397de37b63012b5317198294995b28e8b3b41d3a