Pith. sign in

Paper Citation Record · LEDGER

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2506.01365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01365 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:11.832886Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:09.479147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:51:12.675415Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy26
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef798b6f-3a92-45cd-b13b-cb6e7ab9bd21 · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.275042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.419307Z digest=sha256:94417d597e9bc9092ed52ece41ec86e8cad1bab748ca9cdf915661c3c0a99d2a

Observation 939a4c94-c33e-41bb-a9eb-1d76b740d270 · outbound

This paper cites Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:51:12.791241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.479147Z digest=sha256:b64bbeabd31a2a29e75763e65d40aec91e12c90903880be95754e5549a0153f1

Observation 7d51e26a-34d5-40f0-8512-f376cac3e2b3 · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.248660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.538212Z digest=sha256:19ba8cf6954735da59c3364446492612aec922474b35c2faa70a80fd9d54e591

Observation 37bfde91-3d9f-43eb-8622-27d6fb12c8ff · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.220498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.594732Z digest=sha256:8133d42a1395fa2b4986fb1cfe5d1d87e5503056478810db8e7ea5dd62b46639

Observation 12d97d49-01a5-4d69-842c-ece9c0f60556 · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.192287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.640441Z digest=sha256:5601f343367d6d2cd3d7a2f87fc7a10b9ccace7684d573dc8ddd2b7dca2f74cd

Observation 833c88b7-5636-4adb-96dd-659c7764811d · outbound

This paper cites MFCC vs PTM Features Pre-trained model based features have proven effective for various speech tasks, including V AD.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion MFCC vs PTM Features Pre-trained model based features have proven effective for various speech tasks, including V AD

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:16.162184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.706274Z digest=sha256:de4c022edaf28d766503b635f33996b3563c04de0328d3e77f12623be1e027a0

Observation 49fb3993-86dd-45c4-a184-64e4fb714be2 · outbound

This paper cites Dataset and Evaluation Metrics We conducted all our experiments on three publicly available datasets, i.e., AMI, Callhome, and V oxConverse, to ensure do- main diversity.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Dataset and Evaluation Metrics We conducted all our experiments on three publicly available datasets, i.e., AMI, Callhome, and V oxConverse, to ensure do- main diversity

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:16.136028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.750716Z digest=sha256:26276eedb8d5e9bb1523c37b0638123dab3291732064a29b1650dab3314da42b

Observation 3ef24849-27fb-4d4d-8e30-9ecf004c37ac · outbound

This paper cites A survey of convo- lutional neural networks: analysis, applications, and prospects,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A survey of convo- lutional neural networks: analysis, applications, and prospects,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.879696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.213444Z digest=sha256:3c7f62536274a5ad1eddba4b52526fdaaf6b29a97549df81b2f0e9e7a568c40a

Observation d398ae63-8724-4b6a-9f6b-eb80523e5a3e · outbound

This paper cites MFCC vs PTM Features First 3 columns in Table 1 shows the performance of V AD with individual features in terms of DER, FAR and MR.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion MFCC vs PTM Features First 3 columns in Table 1 shows the performance of V AD with individual features in terms of DER, FAR and MR

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.789229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.860129Z digest=sha256:5074aec3cedd29601515f526e5664c33b1d06b90986e3875b1a4d45c549bc520

Observation eec9b035-31f0-4487-b22d-453ef6f20906 · outbound

This paper cites Our experiments show that simple fusion methods like addition and concatenation consistently outperform the more complex cross-attention mechanism.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Our experiments show that simple fusion methods like addition and concatenation consistently outperform the more complex cross-attention mechanism

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.646275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.900221Z digest=sha256:e32dc5d3d49bc48aaedf5eef56731b1adf5ad76a96571b83439fb37ed222918f

Observation b7cf4512-d6fe-4bc2-90b6-bc82a25d59b4 · outbound

This paper cites Temporal modeling using di- lated convolution and gating for voice-activity-detection,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Temporal modeling using di- lated convolution and gating for voice-activity-detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.569898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.938166Z digest=sha256:126d4faddfd5b177054afcc685be84af62d24c7c0f7a66ccbe02309699ecd44b

Observation fd8a25d1-d16d-4e32-bb0e-4e3e79db3956 · outbound

This paper cites Wavoice: An mmwave-assisted noise-resistant speech recogni- tion system,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Wavoice: An mmwave-assisted noise-resistant speech recogni- tion system,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.488071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.973317Z digest=sha256:36f4b542fbda50cbcbf4af611ab654767e5ca8134cf684beaec771ca8b1455ae

Observation 8222dd9d-4393-4bb4-a8cb-5ce5f7d8b7f2 · outbound

This paper cites Profile-error-tolerant target-speaker voice activity detection,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Profile-error-tolerant target-speaker voice activity detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.321383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.016834Z digest=sha256:729a2cc60c5985bc885ebdc991ea89d05b4b61af5078c3b3238b07c1444a2f34

Observation 679a867e-043f-416a-8b01-a477f1d0aee5 · outbound

This paper cites Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:12.535791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.058899Z digest=sha256:e72794ada82ab457f3e338aada2544b10fda8fc001aeae01c154473c112ebe60

Observation e10244b7-7dfe-4a9b-b0ec-2a8203743f11 · outbound

This paper cites Unveiling the state-of-the-art: A com- prehensive survey on voice activity detection techniques,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unveiling the state-of-the-art: A com- prehensive survey on voice activity detection techniques,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.185183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.100937Z digest=sha256:94553213322d215e7ed92eb8076de1a379371ba15077570fbe688ad187c35f38

Observation 783b5cb6-c1da-48ec-9972-8751fff5b9c9 · outbound

This paper cites V oice activity detection: Fusion of time and frequency domain features with a svm classifier,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion V oice activity detection: Fusion of time and frequency domain features with a svm classifier,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.036588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.141767Z digest=sha256:8d18b225a2c0408b3c9b8eaf42c2de3f3d7fa86acf7dc7d6ebbae3b3f4287572

Observation 5cc4d5a4-f4aa-4fa3-be5d-d4802a5ec1f3 · outbound

This paper cites Analy- sis of derivative of instantaneous frequency and its application to voice activity detection,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Analy- sis of derivative of instantaneous frequency and its application to voice activity detection,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.979766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.179305Z digest=sha256:8fc8c1718536d991ab5a62f5c28041ca0e9a517e4ef2c69307b3087d0085d1f0

Observation 65f42d73-cb2d-4211-8c63-256f6bbaf333 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Robust speech recognition via large-scale weak supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.788432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.788432Z digest=sha256:58e12d331bf688a02ae7f49c0ca7e5c655d22c2008db8bddc955b874a77a1dca

Observation f763f485-49cf-42f7-932f-47987f604b98 · outbound

This paper cites A review of recurrent neural networks: Lstm cells and network architectures,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A review of recurrent neural networks: Lstm cells and network architectures,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.767898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.265665Z digest=sha256:e972a3ee26fc7e285bc0e1cd5899bcb55eb25e786e47e3392e297743919c96e0

Observation 4266d00f-cbff-4121-94cd-08e422b1343e · outbound

This paper cites A hybrid cnn-bilstm voice activity detector,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A hybrid cnn-bilstm voice activity detector,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.698211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.304986Z digest=sha256:1439e1588ea735ff507371d294b1b7654f04aa013e3dc874760d7cce61037bf2

Observation aa3820dd-fd68-4e94-998e-3ed5b946acf6 · outbound

This paper cites Feature learn- ing with raw-waveform cldnns for voice activity detection.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Feature learn- ing with raw-waveform cldnns for voice activity detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.505567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.357047Z digest=sha256:3e420ea9bc25559a5f6f028cca3d1f23c951d841436139f4645a73f7bdbe1c90

Observation e16e1dc1-7fb9-485b-80d5-66fac811bdc7 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.407683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.407683Z digest=sha256:be1ad406f8f2b2e3fe9538e550a1e41a1ae9d692bab3f25b74c1432185019890

Observation 867f4b01-b5f4-4431-988b-1c5c344019ba · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.456958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.456958Z digest=sha256:4df77a7fc0b89724449e17204fe675b559d46d8c99680ac5b97690afd9eb2d17

Observation 2bec1e50-8035-4bad-a19a-a50b21c65200 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.503910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.503910Z digest=sha256:f0c79bb87e495e1efe982fe2f1a60f8405e8c9a80074753fc84b3b7802392629

Observation e161ca1d-c447-43eb-a71c-267d0ee7f257 · outbound

This paper cites A closer look at wav2vec2 embeddings for on-device single-channel speech en- hancement,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A closer look at wav2vec2 embeddings for on-device single-channel speech en- hancement,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.316717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.585637Z digest=sha256:64eadbb014929770d8ba388560ecd101ea5a317d894690f49e5bc4f0f8a47354

Observation 13345372-e7a2-4b62-a873-1fdfa45bfc5e · outbound

This paper cites Unispeech-sat: Universal speech rep- resentation learning with speaker aware pre-training,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unispeech-sat: Universal speech rep- resentation learning with speaker aware pre-training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.190460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:10.643292Z digest=sha256:f7a45c42c9762b7d4009a53f8b4c7fd733273327adf412633a42303e1812e11e

Observation 775471c7-0d58-4ada-adc7-fc7114dd8980 · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Scaling speech technology to 1,000+ languages,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.710355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.710355Z digest=sha256:acad004e51262f5d8e71c00596a61daa290cb69725c1750ebd6aafb29569270e

Observation b353c641-69c0-46cc-baa6-6a632f777a94 · outbound

This paper cites Spot the conversation: speaker diarisation in the wild.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Spot the conversation: speaker diarisation in the wild

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:12.046250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.586294Z digest=sha256:d667b63477018b58dd3288e70cb17682c3e4d03568ea8e0001280642f983dc1e

Observation 70e2a298-5d4d-4288-8bcd-cfa4935a237b · outbound

This paper cites Base version checkpoints are considered for wav2vec 2.01, Hu- 1https://huggingface.co/facebook/ wav2vec2-base BERT2, WavLM3, UniSpeech 4and Whisper5.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Base version checkpoints are considered for wav2vec 2.01, Hu- 1https://huggingface.co/facebook/ wav2vec2-base BERT2, WavLM3, UniSpeech 4and Whisper5

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.960782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.817109Z digest=sha256:33f31186f8a3b0cc80e6937771c1107dae5153fafab0dcbf03c49b3d421a798e

Observation 53a273fc-365b-4118-93fa-4e73e37f9a8d · outbound

This paper cites Enhancing whisper’s accu- racy and speed for indian languages through prompt-tuning and tokenization,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Enhancing whisper’s accu- racy and speed for indian languages through prompt-tuning and tokenization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.873019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.873019Z digest=sha256:5d35209baf2dfc5fbaf661e97737762d84c04a2ee2207688e33775d35eece884

Observation 06ccb6e3-9ae4-49cb-a111-d5754034baee · outbound

This paper cites Self-supervised speech representation learning: A review,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Self-supervised speech representation learning: A review,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.961151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.961151Z digest=sha256:e9189cf8536220872124887e0b910a9db0d5f3310f6c674622e7de5d65e8937e

Observation 620f9e5c-8bd2-4289-90f8-8bc9f2f8eaeb · outbound

This paper cites Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:11.040230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:11.040230Z digest=sha256:132262aa661445452e36212b1a3723c178ed81d6cad72c332dce73ff4d9771b5

Observation 6c1ee7fe-557f-4c6b-880a-fb3f624a8a3d · outbound

This paper cites Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:12.273028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.101007Z digest=sha256:5430814f688afea97e985a4374f2b4ca6590ba73251382e9fe308c30919fc235

Observation b3150781-14ea-40fe-8a97-74a8b0ef0d2f · outbound

This paper cites A transformer-based voice activity detector,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A transformer-based voice activity detector,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.040176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.169108Z digest=sha256:82e913caaa4597c25ebda9641c92dfd2c6b496b03e08d8ea1cb94a5390d51a79

Observation b382a52b-bbd2-47fe-b586-6774fcc0acad · outbound

This paper cites Multitask detection of speaker changes, overlapping speech and voice activity using wav2vec 2.0,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Multitask detection of speaker changes, overlapping speech and voice activity using wav2vec 2.0,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.742840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.271319Z digest=sha256:9de81e168aea386e09e89a5bf8e01882d5c62661e163d620e88ce30958593dd8

Observation b5687662-d01c-41fe-a795-98a53b41442f · outbound

This paper cites Feature extrac- tion using mfcc,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Feature extrac- tion using mfcc,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.589638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.374716Z digest=sha256:52008a34b4e61ba22a90467f4efdd0085ff6268822073661556f947a80e17172

Observation 4b3137c4-f96b-4e44-ad1f-7748b7e574d7 · outbound

This paper cites Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.459115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.439124Z digest=sha256:f9d2b2cdadc7f67bc0497a9c9796a8a31ac7656608bbf891389af1b3e6b84ddf

Observation d0a65c65-e5e5-4e6c-87d4-a2e1b5cc3b28 · outbound

This paper cites 2000 nist speaker recognition evaluation,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion 2000 nist speaker recognition evaluation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.335994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.507324Z digest=sha256:352a38963a2ed10f240c95dceb63b1bc648b98e577db6c6c2378005003dfcce2

Observation 46a94862-248c-4110-ad59-a7f343668c0a · outbound

This paper cites Pyannote. audio: neural building blocks for speaker diarization,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Pyannote. audio: neural building blocks for speaker diarization,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.227135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.666189Z digest=sha256:b06c4356d69f0a60baa755dde1934e89233c065194f1e4c21eebf1e5c7e61c3d

Observation 98c46501-02bf-4a18-b6c4-d86355140ca6 · outbound

This paper cites End-to-end speaker segmentation for overlap-aware resegmentation.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion End-to-end speaker segmentation for overlap-aware resegmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:11.727872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:11.727872Z digest=sha256:113b252aed3c7c5f58151a1b62659ce979fabb6a083f3278b42a28d02c4c0555

Observation 8372bbce-579c-4296-9b7d-c874a795c96a · outbound

This paper cites Speaker recognition from raw wave- form with sincnet,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Speaker recognition from raw wave- form with sincnet,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.170204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.759953Z digest=sha256:56de6dbc6ab1c6936f4698d552064338279c2d0720a7a62beb620e6e364ec969

Observation a3c4611a-7a41-452d-b56e-3aa0e059262c · outbound

This paper cites rvad: An unsupervised segment- based robust voice activity detection method,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion rvad: An unsupervised segment- based robust voice activity detection method,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.021646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.789628Z digest=sha256:9717663575fc92beeaf92f547090d4bb79fd7f39840289f430a97885306c4225

Observation 49dbc3c4-1c47-4149-b9ef-c03473418d92 · outbound

This paper cites Boosted deep neural networks and multi-resolution cochleagram features for voice activity detec- tion.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Boosted deep neural networks and multi-resolution cochleagram features for voice activity detec- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:12.933562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:11.832886Z digest=sha256:651f5d6d0a180c26ef650b01cf64c67475826572aef6a4a9fb19f2f4261329a9

Pith citing papers

Observation 939a4c94-c33e-41bb-a9eb-1d76b740d270 · inbound

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion cites this paper.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:51:12.791241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:09.479147Z digest=sha256:b64bbeabd31a2a29e75763e65d40aec91e12c90903880be95754e5549a0153f1