Pith. sign in

Paper Citation Record · LEDGER

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2505.13971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13971 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:05.021485Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:00.946933Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T08:09:51.394586Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · outbound

This paper cites The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.946933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.946933Z digest=sha256:82592718de5685fc9d602ed530d95053bedf6abc7f02c85b3cbfdc30bc288326

Observation d182b9a8-a55d-4e53-a58c-b0807a475f2c · outbound

This paper cites Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.405861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.015198Z digest=sha256:267cdcbe8e2074ce39dc506aad582a4c0cab771132bdf5a2e754ae29526ad83c

Observation 4617894b-5ab6-4b5d-b4c4-abceef6f4b8f · outbound

This paper cites who spoke when.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition who spoke when

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.171540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.131980Z digest=sha256:33f6b11b6247365eeb6f0f529e8d5281098743b0f4e3e4795f6648a962ccab58

Observation 63d2ca25-9035-4f8f-9644-6e043c78bd31 · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:10.001010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.253337Z digest=sha256:7b3b13b0a48646e37ad5f891d93d136ade93a20550487381ef67d8b65f095333

Observation af8c5e2b-7ce1-461d-9630-7081a33c5d9e · outbound

This paper cites (2), whereNspk is the total number of speakers in the session.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition (2), whereNspk is the total number of speakers in the session

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.796298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.356018Z digest=sha256:ce1687fea403786c9e69453d74eb42a8f636dc80bdfa9b4b36db362b39d7f1a1

Observation fd03b07a-e841-49f1-965a-560307c032da · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:09.545159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.496388Z digest=sha256:265a62f0b492a05ca5401c8edc526fe43032b5f7ee8a7e87b6d7d13eb3bad4b9

Observation fabb067c-e0b7-4de2-9160-4bc51668a496 · outbound

This paper cites A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.281598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.616838Z digest=sha256:23c44b07f0c9852c67e0f6ec9b431305712a61a65c176270839e4ca1dc600cab

Observation 095c7367-1392-4e02-a4bd-1df23564b5f2 · outbound

This paper cites an unresolved cited work.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:09.111934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.738485Z digest=sha256:90fd89f044a1778c892328ddb6a4f08d3b912cececc3b2009328bd6799110f44

Observation 1ee43906-d77d-4382-a9c1-11056bb6132d · outbound

This paper cites The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.828947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.825645Z digest=sha256:66c397cb68b5449d129c003ff35bfc61ac62cdc76a9591dd5b732aa688df71aa

Observation 94b03a91-ce32-48d7-8791-45fde30cb4a5 · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.634867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:01.970675Z digest=sha256:f5f053863f2d179afd1cbc200ef88a90e7ee7931ec82126c1fe1278e5d3e0838

Observation 474dda51-9528-4bbb-86e4-476a84c651c5 · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.079339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.079339Z digest=sha256:99a57861d16e7a2964619934e7e15029e2a842b52f1f0c662600a9acbb3caf8a

Observation 24dda821-8acf-4170-af19-ffb35fec1c65 · outbound

This paper cites Continuous speech separation: Dataset and analysis,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Continuous speech separation: Dataset and analysis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.436561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:02.211446Z digest=sha256:9fb044fd97e8fd97ad24c6ab2eabb9fbbfd0969a186ba6e9c9685ca181ec7844

Observation e9c3e128-15c4-4b6e-ac08-f377e39382e2 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.305820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.305820Z digest=sha256:d0f81ae09c03bd19f250f9bb4a9150ab68d52fb1e55ae546cf2fb36b173814a0

Observation ef7adc73-dd1c-4ab8-b3d1-cf8b99dd345f · outbound

This paper cites CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.186566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:02.459989Z digest=sha256:5ac7ef1f53dca41aecc0cb5373f64140959fbc1774175237f2d3c77a75a763c9

Observation 5b477c2d-8e4a-450f-bd4f-280b2da4535e · outbound

This paper cites The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.989009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:02.614902Z digest=sha256:f63af57ed5b3fbede470ef71cb2f04fd398d75c085855a5e471464dd3992907f

Observation bea85d7f-dd75-41a5-bfda-ee982ef8a884 · outbound

This paper cites The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.797124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:02.797124Z digest=sha256:2670a1a0e62e0a197289920eb7ca4d70ea66f3ef60ba201f14ea4302c74c7325

Observation 9c480d17-e3da-41ee-96c7-d54cb0103dbc · outbound

This paper cites The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.764255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:02.876803Z digest=sha256:b877918e642b547a47ab57965bb5e63b3f87e8cccd952c24b5931197783866ab

Observation c4e4131e-df16-4fd3-abab-13d0adc854ab · outbound

This paper cites MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.014747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.014747Z digest=sha256:c90a3618539ee1fcfe8b66dd8e42aa0e0cb5e35ffa8dab8e9e4036c1bf8a1b3f

Observation b75b325c-4bf8-4d38-a95a-702aac5698b2 · outbound

This paper cites End-to-end audio-visual neural speaker diarization,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition End-to-end audio-visual neural speaker diarization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.550360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:03.076635Z digest=sha256:3a4ecabf0eab5585235325469f4a75d4b79a24000e5cfede2682576aaab96426

Observation 75b3023e-f670-4324-8b08-aae5b29f9a25 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.146074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.146074Z digest=sha256:9f0eb08137a29a54c1f7a8ca2efa9c4a036fa7d9d74c9b4879776528714d3a9b

Observation 26e1751a-a452-4788-af8f-6fa7d246e9df · outbound

This paper cites Dover-lap: A method for com- bining overlap-aware diarization outputs,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Dover-lap: A method for com- bining overlap-aware diarization outputs,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.211149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.211149Z digest=sha256:0fe280b0822ea182abbb470e5cb0a5145dc165cf849fa801d81a26dc805bfb9f

Observation 5d52c1c3-c08a-496c-be86-9f1f01c3556f · outbound

This paper cites The rich transcription 2006 spring meeting recognition evaluation,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The rich transcription 2006 spring meeting recognition evaluation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:07.172677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:03.345208Z digest=sha256:40668896688301c526e780f9b1728bcbfcfbbea354bdc06551e0539ee0159b46

Observation d7c11625-c882-412a-bdda-0e38ad429a5e · outbound

This paper cites Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.921969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:03.405144Z digest=sha256:4cff13b84a977da39c7031db01fa6942e5a913550363dc5dfb6330c35cbc652b

Observation 0458f9e8-8c9c-4667-956e-6af705fe8edf · outbound

This paper cites GPU-accelerated Guided Source Separation for Meeting Transcription.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition GPU-accelerated Guided Source Separation for Meeting Transcription

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.476909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.476909Z digest=sha256:9408ed007761ddde2ef29a521ad70408bf7b9bbd9711ceafd66c176879814d93

Observation 92676c42-0012-47ee-9ef8-aa8e151969eb · outbound

This paper cites The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.650135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:03.600303Z digest=sha256:f5c8efb809a90ccce87178bc30dadcea54c944ea772695fd8ec7c1cf9fcc4b8f

Observation 90723b50-6955-4e15-aa8a-5229bc7d5a0b · outbound

This paper cites VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.767045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.767045Z digest=sha256:f2b84ce25bc59f9cbeae099e4dd82e2e7f7d1c7b8d6659eb1507ff46c3d8fa58

Observation dc3d2ccd-9969-41dc-ac57-55fd0e06b23b · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition VoxCeleb2: Deep Speaker Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.878807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.878807Z digest=sha256:320579e85f236676cb980a0c3e6590b124125c086ca72e0b9e1ec39160d85269

Observation c7e3d097-71db-4f79-bc87-2ea03ec6a50e · outbound

This paper cites Kespeech: An open source speech dataset of mandarin and its eight subdialects,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Kespeech: An open source speech dataset of mandarin and its eight subdialects,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:06.293656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:03.938904Z digest=sha256:1135854158bee1ce65e6fbbbfc6590dd200be6d2a4caa8e26dbcb11bcfc5f468

Observation 5ad194ca-b72e-4f4d-9468-e213687d7f26 · outbound

This paper cites 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.029122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.029122Z digest=sha256:f93d6ffa8718f178ba37739d5a871a15b53004ba565036f0bf4db12e420a67d8

Observation 2fbc7da8-9942-4d95-b08a-e464e0796a9c · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.140000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.140000Z digest=sha256:ff57cd90cc39e22462ae36bc88d123612fffa87cca1e07ead4934d287eb60118

Observation f9b28de0-ed4d-4341-9501-1ae12bebade2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.219246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.219246Z digest=sha256:13abab6f621351805296668cad3e00af7b2dc76f7f2152d018b1ebdc10f9e458

Observation 498eb297-5d96-4a18-9d4f-b1523a25db7b · outbound

This paper cites The ami meeting corpus,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The ami meeting corpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.308158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.308158Z digest=sha256:6a1163462c136f8a136e73d9b405ecb8444f26802ea1159ea42ee64d6598d94c

Observation 49055ca4-5d30-4e6b-b90e-6adb66cce4ce · outbound

This paper cites Lip-reading with densely connected temporal convolutional networks,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Lip-reading with densely connected temporal convolutional networks,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.390574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.390574Z digest=sha256:5484fb5107eb45ed72634b93d3081a32b9d9ab93cc55e314148d32eb94329c8a

Observation ca19d67a-28fd-4b16-9564-edaf8e684aaf · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Robust speech recognition via large-scale weak supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.485577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.485577Z digest=sha256:3c956db6285eb074d42e23d06323337d21c78895b9e2ecfb74dc5e39eb5b5840

Observation 6dbbfdd8-4a91-4c43-be47-b2a52ce12a48 · outbound

This paper cites Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.746106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:04.559893Z digest=sha256:ee4260b8113ab190567b5e195f798d16c573bf3793463bb7ad1d150519e137bb

Observation a70e8cee-0bec-4546-b2c3-3b0c3e71d1fc · outbound

This paper cites Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.678989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.678989Z digest=sha256:a73ea487e8a842ae7985dfc80865bf652a0b682be4d3e6089dea7edc87aeeb05

Observation 909b7d81-34cb-44bd-b0dd-6df558509c42 · outbound

This paper cites Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.846296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.846296Z digest=sha256:7fa44fa62341881640d82b48a85719043f25f9e2c388571b9f4660b2cd32ded2

Observation 1494347b-02b1-43d7-9318-d5ff6fcb3771 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.928529Z digest=sha256:675e91c1a43837079138ed5cc0be18a6ba0ab8d0349444d4fe454a60d08d9e8e

Observation fc4ac28f-13cd-4e27-8b80-0d5abb96a892 · outbound

This paper cites Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:05.514736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:43:05.021485Z digest=sha256:63b0453576c65ddc9071b92497f73beb4c7f73b35433ae9fe505205da9dda428

Pith citing papers

Observation d516a4d6-1b5c-43fa-bd6f-344aae0d6d84 · inbound

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition cites this paper.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.946933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.946933Z digest=sha256:82592718de5685fc9d602ed530d95053bedf6abc7f02c85b3cbfdc30bc288326

Observation 02c18a0e-adb3-4d68-9574-028150db69db · inbound

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge cites this paper.

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:26.572189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:26.572189Z digest=sha256:44ff2fe8d8d5ce3bf8e51979487b0df908445a79090299a8ae923e9004dca3c4

Observation dca6c443-3407-4b0b-92fb-977eb884cbde · inbound

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset cites this paper.

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:21:56.777201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:21:56.777201Z digest=sha256:226c3226320890e3aafcf1bfb088014aee2acdcfce0b8b864d98347d23e37a93

Observation 06774b7d-55b1-42a7-abae-e6f845a405cf · inbound

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs cites this paper.

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:09:51.398870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T08:06:45.477344Z digest=sha256:f4b3b6a883a44222a1e66732c132cdaeb452d1ce3be6ac01e790de8bc9badb7e

Observation 038ca04a-f233-43de-a3f7-8b08fa5f8d8a · inbound

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models cites this paper.

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:12.615570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T09:18:53.285951Z digest=sha256:a0bfdc092a24eeda75d8bfba0a5ad43cfc57114282c089db18fb3062cacc59ca

Observation 38be85af-3c4d-473d-960a-35da92a73616 · inbound

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings cites this paper.

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:06:07.806541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:06:07.806541Z digest=sha256:77edec11ac1e5c838b29c3af2d1267df5d23b91fdeb2de34bb8ea6e8c39ccab3