Pith. sign in

Paper Citation Record · LEDGER

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2506.14427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14427 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:21:58.204834Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T15:07:52.715880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.188094Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy43
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3f82b19-44d6-435f-9c02-be0f261f38bb · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A review of speaker diarization: Recent advances with deep learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.659626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:50.470683Z digest=sha256:c905624104291abb187f4f3c002c45ed93f59807f53a385aa1f4dd7fc986a961

Observation e13f9fe6-6fd1-4b5a-961a-4aeb8dd64713 · outbound

This paper cites Speaker diarization: A review of recent research,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization: A review of recent research,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.458552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:50.640545Z digest=sha256:c265b102366dc5c8cb3c8205a2cc0b88cc80cb4fad81e3c441d3398337cd3336

Observation ec32f24c-0d18-4624-b2e7-4f1c8a84a765 · outbound

This paper cites Speaker diarization with lstm,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization with lstm,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:13.238820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:50.787115Z digest=sha256:974c380b45ff83103160431244edd2e54995097b1daf9122611a0e5ceff77dc7

Observation 73ef4eec-b169-4a19-99fd-25a05bff166e · outbound

This paper cites Front- end factor analysis for speaker verification,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Front- end factor analysis for speaker verification,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:50.865767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:50.865767Z digest=sha256:b3a0d13e9ac77ce8e1e1f4633c9f783a4071a24771543d7b434ef490f97c6700

Observation 9c5e0ead-22dd-4f87-8f96-979565dfbea6 · outbound

This paper cites X- vectors: Robust dnn embeddings for speaker recognition,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset X- vectors: Robust dnn embeddings for speaker recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.974964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:51.002901Z digest=sha256:c9737ddb1cce465fde3ba9be122264157c783d8214968fb0cfed77f029e9ce1d

Observation 56007a15-7a95-4dad-9254-627bcce8099a · outbound

This paper cites Developing on-line speaker diarization system.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Developing on-line speaker diarization system

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.703109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:51.144910Z digest=sha256:74dcc7f6ea9874d4eab8b79bf96cf201e083f4eafd6b4d98df6659ce1fdb561f

Observation 1efab697-e996-43c9-b157-0884261e392a · outbound

This paper cites A study of the cosine distance-based mean shift for telephone speech diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A study of the cosine distance-based mean shift for telephone speech diarization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:12.304830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:51.314753Z digest=sha256:352abd385386b9ca311b1b236a459a1fbb5d3dade9d1c1471c22d31ababff246

Observation 09e744c1-aa83-4cfa-bc3a-41559e4f6828 · outbound

This paper cites Speaker diarization with lstm,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Speaker diarization with lstm,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:11.883861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:51.517631Z digest=sha256:bc7bc0e947ff4b514884283ebc967b93a61736cdd33601d4a119cec19993be77

Observation fde14acd-eef8-4b76-a3dd-e02491bab023 · outbound

This paper cites A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization system

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:11.374824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:51.731444Z digest=sha256:46cb24751e3f92cf0b098b8ca8051c0ce2cb38250bfd99b288f3d51e675014ac

Observation ad39687d-62ea-4546-9eb8-9db7c3a523f3 · outbound

This paper cites Characterizing performance of speaker diarization systems on far-field speech using standard methods,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Characterizing performance of speaker diarization systems on far-field speech using standard methods,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.884984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:51.890542Z digest=sha256:9bc97adda315993fdbf3e7f54f5724e4ac1ed1f3f285b3da1ad83e4fce1e9556

Observation f87ca2c2-33fb-4503-8ebe-fe82f85839e9 · outbound

This paper cites Discriminative neural clustering for speaker diarisation,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Discriminative neural clustering for speaker diarisation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.494758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:52.094815Z digest=sha256:a4fdebe8a7e7f40545e45ed4558797d27612b4cd2995d5e73d4a72114919bf9e

Observation 1799a67a-a6c6-4938-a63d-5af4012ce448 · outbound

This paper cites Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implemen- tation and analysis on standard tasks,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implemen- tation and analysis on standard tasks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:10.129979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:52.305116Z digest=sha256:2270898999198239fa953c68d5964599b7ac248bba215f99dc5e6da260615592

Observation 6c09ff5f-6367-4f97-b3af-add1a7bd70f9 · outbound

This paper cites End-to-End Neural Speaker Diarization with Permutation-Free Objec- tives,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-End Neural Speaker Diarization with Permutation-Free Objec- tives,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.865817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:52.444837Z digest=sha256:79700631e16af3f5efaf2b709dbf3c73e847fab53f3aa59d590c651dec40875b

Observation 6962df6d-2f1d-46a9-b7ff-7e1488b08cd9 · outbound

This paper cites End-to-end neural speaker diarization with self-attention,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-end neural speaker diarization with self-attention,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:52.593916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:52.593916Z digest=sha256:de710c2a3d3bbbec8214f849b0d659acd0d637e1113feedb127a25c2f5fc2520

Observation 7a66dfec-ffab-4768-bbe7-c47f87b803ec · outbound

This paper cites Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.614817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:52.694892Z digest=sha256:f07fd2c9d30b47846a34b0a939b4e21a4e5ee69b287128638ed4ab236e396890

Observation 91d2595b-6fe6-4cd8-ac5e-ddf332add929 · outbound

This paper cites Incorporating end-to-end framework into target- speaker voice activity detection,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Incorporating end-to-end framework into target- speaker voice activity detection,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:09.328883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:52.834754Z digest=sha256:4d7a4fccb7e69e320748a4e6ae9b3b32ebf4a8344e7e27c96c2341b42bb2ab40

Observation 48b1ea25-89ef-42ed-8019-a486ea98782c · outbound

This paper cites Target-Speaker V oice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Target-Speaker V oice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:08.528593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:52.935871Z digest=sha256:737506ee5b90c7747389b8ad77b99460529e74c272b59eb02e52d46fa3a85234

Observation e72957de-844b-4af2-b005-b6d83bdb7240 · outbound

This paper cites Ansd-ma-mse: Adaptive neural speaker diarization using memory-aware multi-speaker embed- ding,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Ansd-ma-mse: Adaptive neural speaker diarization using memory-aware multi-speaker embed- ding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:08.157893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:53.134746Z digest=sha256:c1d46ad6df9bd5a8b65523a1f0d1f6943c08f67d88153f402246f8d0c0c72daf

Observation 2716b444-1e63-4fc5-9248-731d85914f91 · outbound

This paper cites The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:53.248557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:53.248557Z digest=sha256:0c0d5a668c1f682db5322cb64d8bf9fe23c6f9fe5e939101e5e5ea880dd7c7ac

Observation 8ff547e6-29f0-42fd-a60f-6f86231dddc2 · outbound

This paper cites The ustc-nercslip systems for chime-7 challenge,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The ustc-nercslip systems for chime-7 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.773098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:53.427755Z digest=sha256:c3e5e683b7432925563dc80243569580537a105a0d95fcf67f4a85e1332ad488

Observation bca27e2b-c9db-4d31-a9a3-d0314d2ed985 · outbound

This paper cites Audio-visual speaker diarization based on spatiotemporal bayesian fusion,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Audio-visual speaker diarization based on spatiotemporal bayesian fusion,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.537163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:53.571500Z digest=sha256:8e19aa7ccc79e8ae01cc3a857b0f8aad0898e2662c12cc29e4d1421941af4042

Observation c773421d-5f15-4cde-850b-8864ec10fbf1 · outbound

This paper cites Quantitative association of vocal-tract and facial behavior,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Quantitative association of vocal-tract and facial behavior,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.309079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:53.778190Z digest=sha256:4ae135be59c42decf6340845b013c1c1f59f4f0fa0c4f7e1c49322147b075e4f

Observation e43a8855-a018-4c74-aa81-c275a07f6cf7 · outbound

This paper cites Multimodal speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Multimodal speaker diarization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:07.064881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:53.934772Z digest=sha256:b3e931859a62f34767bdc3a20f26f6f45cc71435c13002d9f2642a9795d3bf1d

Observation ce9ae35e-a165-40e7-985d-40c08a73cb89 · outbound

This paper cites Who said that?: Audio-visual speaker diarisation of real-world meetings,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Who said that?: Audio-visual speaker diarisation of real-world meetings,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:06.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:54.094834Z digest=sha256:f7c85e4aa2f06aecdf98d71c956def1466d046492c21011b4abf21926a7a89dc

Observation 2e95f8fa-c44d-44f1-87d8-08440b71e7b2 · outbound

This paper cites Spot the conversation: Speaker diarisation in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Spot the conversation: Speaker diarisation in the wild,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:06.436582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:54.254227Z digest=sha256:43e9a9b370f9d46f22ea0c53f8cd85878cb8bd705faca2df62434918a501f17d

Observation 428b31c9-85f8-4f23-9c29-9dac8e0da984 · outbound

This paper cites End-to-end audio-visual neural speaker diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset End-to-end audio-visual neural speaker diarization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.908267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:54.571101Z digest=sha256:35e847f5361a69ca6e9a3138e1035273af47c791a5525b41b5bf886bbd55fc26

Observation c140ea81-7ae9-40be-b600-f63c29793861 · outbound

This paper cites Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:59.726773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:54.869454Z digest=sha256:94da1331c3d970f2a6bdb10dcba4088f41bcf29721a3ad42f6b41e381f5ccfd7

Observation 0fa53158-fd14-451c-907c-60b3a82249c7 · outbound

This paper cites Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:59.396178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.024954Z digest=sha256:3c3ef36025599e3d6f69caae155841190de5d87241ddff2d4a048751fd24ae08

Observation 2ffe891a-1a73-45a7-907e-004876e43b21 · outbound

This paper cites Semi-supervised multi-channel speaker diarization with cross- channel attention,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Semi-supervised multi-channel speaker diarization with cross- channel attention,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.602331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.124886Z digest=sha256:2b841d7d8d966fc80300b5cd386eb758fe331d863047480cf0f8b16f1bbbf7df

Observation e1c36b09-1a9e-4a1b-8abf-6cfe9b2aafed · outbound

This paper cites Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.349237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.207393Z digest=sha256:0b7f079e8ec4b224a50260c4c9bcd4568d55eba0bc4419394f82be0e35a9f53a

Observation 026db9a8-134e-4b29-b845-835258614c2c · outbound

This paper cites Ava-avd: Audio-visual speaker diarization in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Ava-avd: Audio-visual speaker diarization in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:05.115533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.324692Z digest=sha256:5578d524ba4218554aca9da0308dac822dd14faef2426b7c02873374fd684f2e

Observation 05fbe6af-a073-45db-bb76-3551a48e1ab3 · outbound

This paper cites Semi-supervised training with pseudo-labeling for end- to-end neural diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Semi-supervised training with pseudo-labeling for end- to-end neural diarization,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:04.756713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.483931Z digest=sha256:45a905395f158fbb37fafcdc81720cd056356ce14d314b090e54b8423ed8f708

Observation f4fd928c-abbd-4ec0-8ea1-5b27ae7b4a26 · outbound

This paper cites Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:04.309737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.624749Z digest=sha256:11b55a868a7a5fc7228849d304aed44871a9bd6dc86458f2b86dfad1d6145af8

Observation ac8da1ae-9ff5-4ccc-b653-e77e25873ae7 · outbound

This paper cites The multimodal information based speech processing (misp) 2022 challenge: Audio- visual diarization and recognition,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The multimodal information based speech processing (misp) 2022 challenge: Audio- visual diarization and recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.958320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.784828Z digest=sha256:801b273ca97c4445cd547b0b165eaf1857c7286082aa278fc25cc5fa4674d076

Observation 284f3f73-a188-42a9-86b4-331cfea318fa · outbound

This paper cites Notsofar-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Notsofar-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.664897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:55.902011Z digest=sha256:9d8c60d7399ed18f4f40aa7696d71523a9fdc22f3d19dcbdd772d2ded241f346

Observation cedac64d-e72b-4cea-976f-faf267d75d7e · outbound

This paper cites The nist speaker recognition evaluations: 1996-2001.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The nist speaker recognition evaluations: 1996-2001

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.410058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.047128Z digest=sha256:10f359b44222d2e555548360888562a301a07d0d6e8f89de366e6605eae6a187

Observation 6dad5101-8ed9-4e39-b33c-813a367f287b · outbound

This paper cites First dihard challenge evaluation plan,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset First dihard challenge evaluation plan,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:03.186514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.162241Z digest=sha256:c49cd0f9d64d7a335c00f7e469c459d375a7f60f5d295cea6cabd3d0ab450a8c

Observation 68cf3cf7-1053-4eba-8bcb-860721340f5e · outbound

This paper cites The Second DIHARD Diarization Challenge: Dataset, task, and baselines.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Second DIHARD Diarization Challenge: Dataset, task, and baselines

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:21:58.984753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.288230Z digest=sha256:f766287667dc8c1242cf1ea61395adc6a3b63601db2f19415e2afa89728a4a24

Observation 7b1e1bb5-efde-4eee-9410-07ad6e9434eb · outbound

This paper cites The Third DIHARD Diarization Challenge.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Third DIHARD Diarization Challenge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.433967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.433967Z digest=sha256:62bce78c38e05da0ff188edfa7e8ebbbf197b9d3a6c4ab6df5636a577dce1a54

Observation 5daffbd3-c75b-4dee-adf6-aea506583848 · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.864740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.562953Z digest=sha256:452adf67d1578874c50215e9b6e3dd14a84ff86578a05afc82b279b64927940a

Observation f88caf64-7715-4f24-94b6-681e862c35be · outbound

This paper cites The ami meeting corpus,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The ami meeting corpus,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.667720Z digest=sha256:7ae4b70ac2cd55d73215fd3d22d99a73ae6a230aec153f30b589557f8d65e2c7

Observation dca6c443-3407-4b0b-92fb-977eb884cbde · outbound

This paper cites The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.777201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.777201Z digest=sha256:329d269e7d5b38f3805c7852c11fda2f3e2f42724943121ade75142fa723130c

Observation 5438a602-d4a0-4270-b8f1-af5dabe50ec5 · outbound

This paper cites Msdwild: Multi-modal speaker diarization dataset in the wild.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Msdwild: Multi-modal speaker diarization dataset in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.314881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.893682Z digest=sha256:e0814e9580df76f76e1aef88cc59955e1ef9e9a1a0af2500895f25dbc5520eb4

Observation 17e3e495-23e7-4b84-b33e-6598d3733f73 · outbound

This paper cites An approach to scene change detection,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset An approach to scene change detection,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:02.056694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:56.984747Z digest=sha256:c0235dbfa60120bb08070f0189a9b5352e6748ee8adcd7c88e2f50ff3fb03a88

Observation d947b473-73bc-4c55-b50d-545945fa1c29 · outbound

This paper cites Dnsmos p. 835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Dnsmos p. 835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.804829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:57.111910Z digest=sha256:5a6caad136fb2caee167ff64d01ae272faabaa6aa2e5b940b2945c697f14d25e

Observation b6875ee2-91e5-4828-9401-ffd2f0c4e60b · outbound

This paper cites Md-vqa: Multi-dimensional quality assessment for ugc live videos,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Md-vqa: Multi-dimensional quality assessment for ugc live videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.506885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:57.220765Z digest=sha256:2960c5fd0ca8134dd80027585a538add7d9a7f337815b59994144c560068fc43

Observation c6a24fa6-3ee0-42bd-a628-9d00325f3f8b · outbound

This paper cites Out of time: automated lip sync in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Out of time: automated lip sync in the wild,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:01.255798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:57.380226Z digest=sha256:cfe1da1ecaa6bd90ce6707891778f42ae6fbd3570a4f233c2b5137b210408ccb

Observation 8a3154c0-925e-43c5-bf15-519b2d0f01ca · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Retinaface: Single-shot multi-level face localisation in the wild,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.979743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:57.559593Z digest=sha256:fab3dcfd0849c5c27d9dfc47ad3ee1b17f8e68e26dd591c54e04b685e7569733

Observation 6c5c5a94-fc8e-4f58-be3c-5990295ce602 · outbound

This paper cites Simple online and realtime tracking with a deep association metric,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Simple online and realtime tracking with a deep association metric,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.738802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:57.702925Z digest=sha256:0773eb20fbe20e8efd05f34bbcb3d1c96e3220afbd9d605c7a9405d9aeba08cf

Observation a6efce1f-5493-45e2-acdb-574c78d25f91 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:57.834750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:57.834750Z digest=sha256:2967b58bf14d683a39d0af1838e79c40e38e88359e19da4e46b398df6e52bcfa

Observation 7e4dfd2a-1d75-4bae-a506-09f2c1823712 · outbound

This paper cites 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.464756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:58.044836Z digest=sha256:4b0f252f113d48c9d5d5eb647280b01a8c1a6c5743a2534bb937c092dc457b5b

Observation 10058e43-5b10-47f4-b6a0-c3e06c096d82 · outbound

This paper cites Dover-lap: A method for combining overlap- aware diarization outputs,.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset Dover-lap: A method for combining overlap- aware diarization outputs,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:22:00.209099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:21:58.204834Z digest=sha256:70adbdb997f5eaf7b830e108bcd8d2069f44e88b358a10925936f24d29d839b0

Pith citing papers

Observation 627f9364-7480-480c-a8e3-e381cc071047 · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.189364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:15964c4c8caa399598e0105d65dce30b5f75dada2fcb2d8911ff6d9b20d8b2dd