Pith. sign in

Paper Citation Record · LEDGER

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

As of 9 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2502.10447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10447 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:46:12.269019Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:39:34.251185Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb00bbcc-96c2-42e8-b779-f4adb89613c5 · outbound

This paper cites write newline.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.941722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.941722Z digest=sha256:50b5b5162a4b290201d7638ef74951ecf50d66c50d087f22bddff5d79f281a7f

Observation f8a90e81-c9e9-459c-87f8-006a926ba1dc · outbound

This paper cites GPT-4 Technical Report.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.947979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.947979Z digest=sha256:8dc731710206e2e4c88f70699b152469a5c0f0c73732b6f31ec2a156d0e4c1fb

Observation 78d4d530-6232-44cd-a438-e5cf4c89d29e · outbound

This paper cites S., Senior, A., Vinyals, O., and Zisserman, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Senior, A., Vinyals, O., and Zisserman, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.952950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.952950Z digest=sha256:5b58f932a3a7befb76dccf0bfcbeafa9fd980c3867d3eec373646f24c0254624

Observation 09235db7-3652-4033-92ad-731095d6de81 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.957891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.957891Z digest=sha256:99b04d3db7f5d75b56c4cb76eb3b7af76d1ac75f5e828b156268addc3b3c0bad

Observation 2e695b4c-6200-41a6-b137-6a12fb193dd7 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.963258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.963258Z digest=sha256:37c5bf31f22259b2fa58130174c7f38583cc1c8af4eef6961f7435788076c356

Observation 6caeebfe-07b7-4133-9379-0b036559bba2 · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.967831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.967831Z digest=sha256:b6ba37b88dbbb9b5de62e0531c30cc5ceba833caf54f861708f2608b56ebcbec

Observation 0d56858b-ac09-4ac0-ad31-d73f45facc6c · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.105633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.972414Z digest=sha256:838f621c6bde0b8d697b9f4c5b4aabfb6e801bb3db65c504ae2bcc8260d691fc

Observation 3dbff9a2-7450-40d3-b556-096fbdcda2bd · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.977100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.977100Z digest=sha256:7545245a2e3ac17ca0f345e233591837dce443a774740cb0d8cafb250eb81140

Observation 62f50549-e5f2-4082-84c2-aa4a81d73ecb · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.981186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.981186Z digest=sha256:42c89243ec8dcdec3e240762938d1b48f70bd06d6eee8fae1ccc4bb34b444d20

Observation a4b7e83e-c38c-4e3a-8a4a-ccc9451e40c4 · outbound

This paper cites and Timofte, R.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Timofte, R

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.089504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.985691Z digest=sha256:eed68773db22aceeff3ef44225c4dcb5d2487ce50e1c7223a211beeef7e75a29

Observation 9d21efac-daea-48f0-a0b4-23f51b2e50e2 · outbound

This paper cites Large Language Models are Strong Audio-Visual Speech Recognition Learners.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Large Language Models are Strong Audio-Visual Speech Recognition Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:11.989857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:11.989857Z digest=sha256:f9a9fa40d3835e7a5ff8032ebdb191c1f232872f4162879c94275a21118e3f9e

Observation cd483a05-4471-44fb-91fb-0d045342d569 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:13.080054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.994730Z digest=sha256:45df802e31b961629742cfb053f70fa5ecc05b0ee62b329a0d2686e0e0a49f49

Observation 94c28bb5-c138-44f7-9c7c-6d1910553bdd · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.070272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:11.998955Z digest=sha256:136e80ec6c6f7330aab8203c4be949f75687541a25382784055c4c9a738088da

Observation 0313e3cd-1a74-4706-b120-33ba3729349d · outbound

This paper cites Mixtures of experts for audio-visual learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtures of experts for audio-visual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.060255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.003079Z digest=sha256:eb4dac6d607f2b493e3d9b822fb1eba30e2a70765c099ccd56890ce8a7eebbdd

Observation d3bee3c5-68f5-4dab-b490-218a46cd56b7 · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised learning with random-projection quantizer for speech recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.050389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.007423Z digest=sha256:f9094d07ccaa09df2461524d3c93e70f4e2719ccca4ec705ca44468c94172b3d

Observation 8b5cf09a-2a9a-4c10-b50a-b937abbd56ee · outbound

This paper cites J., Kim, M., and Ro, Y.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition J., Kim, M., and Ro, Y

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.040422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.011353Z digest=sha256:f0e18b11e09dc78ff790811a96192ffc8e7567bc7a5699add10ba62bb1da5497

Observation 10622f6a-273b-455e-9a13-874523981cac · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition S., Nagrani, A., and Zisserman, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.030425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.015208Z digest=sha256:71e4ee05cfea2696c699105ddd319e842e0ff37bab5dfca58a77ebf609928c07

Observation 5e64d104-7cd0-4ae7-8d51-06a432a874d7 · outbound

This paper cites Unified scaling laws for routed language models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified scaling laws for routed language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.019361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.019021Z digest=sha256:226c1d447f3fac7ebaa7fdcdb32e0e061a721ca1c7753e6627aec331b6b15d51

Observation 5dd41cc2-e5c3-4328-aee0-e1070ad4bbc5 · outbound

This paper cites Stablemoe: Stable routing strategy for mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Stablemoe: Stable routing strategy for mixture of experts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:13.008786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.022925Z digest=sha256:d87fef225dbf6de582603ab67ae8eb2a981ac67ff136cda8b04426d67b5e7646

Observation 0405de2f-f8a3-485f-94b7-c584ad2fa1ec · outbound

This paper cites A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.998388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.027308Z digest=sha256:e689dbaf3a47cedce438aa0207554002762f82df5cfb62f9eaf0177747d48ac0

Observation 5fa66c23-3834-481f-b695-ae71b99e2503 · outbound

This paper cites and Luettin, J.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Luettin, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.987504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.031328Z digest=sha256:78e7790f7c975dd184a1a69e06a148d545f51e9c91d55c2565b8f9f9d37e732a

Observation 54992c07-7d83-4e18-8e1b-8fd801d9fcd0 · outbound

This paper cites W., and Matt, P.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Matt, P

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.977185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.035295Z digest=sha256:684445dd042131ce0765f62bf76c090db9ca050c04ca7d6017a8cefa1e99a8d9

Observation eb738fe9-acd8-45a5-8078-6693ed7466bd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.039182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.039182Z digest=sha256:10e78ad1a512a0bcc93e09685686c0f414155e558da85ea928e4bab82fe03a59

Observation b727f062-e417-4fa9-9016-14e32e64f9ff · outbound

This paper cites Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.959944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.043299Z digest=sha256:fd0346e968dbae37033ec95ca36c4cfac9969b8555e62346ba028f21c0d6e226

Observation 96a8fece-b9ff-4f10-ac66-9350ea481c97 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Conformer: Convolution-augmented transformer for speech recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.949631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.047254Z digest=sha256:39ed4f25e5bd7cd53cc6ebd16e056d0a6c0edf359112eeb71a1442b339a6130c

Observation 0386b6f4-63c0-4444-8e36-a72ec75d3e59 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.051209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.051209Z digest=sha256:3de3f2b55420ee24ee3056ecda75fff38e4782f0e23465047d30321525e7cddb

Observation 6f54ce98-8b96-4709-a4c5-b81514cd5cea · outbound

This paper cites Jointly learning visual and auditory speech representations from raw data.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Jointly learning visual and auditory speech representations from raw data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.938761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.055256Z digest=sha256:65baf7ab8b047865071f8f9d45b33d4bb0b46ecf9ee64040122c4f85ff6e1562

Observation 55a73f4d-4585-42da-9190-fea994ad50c4 · outbound

This paper cites Braven: Improving self-supervised pre-training for visual and auditory speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Braven: Improving self-supervised pre-training for visual and auditory speech recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.927405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.059701Z digest=sha256:9bfd2989d63bd1f48605cb55fe625f37be83ca313ffb7b8bacde6c01dd8b6342

Observation d65721a7-5e49-4b3c-9e82-243a295f819e · outbound

This paper cites XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition XLAVS - R : Cross-lingual audio-visual speech representation learning for noise-robust speech perception

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.917268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.063242Z digest=sha256:b3cc241a94991088f35d529752965cb041dce7925bc1ca816c0a61db12ee5ceb

Observation 9c4f2e9f-e739-4545-a79d-fa40926bb734 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.906730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.067010Z digest=sha256:0f1abaf8b946c5b12c8c48a3169ddbc0447b5f31563d889e0c174c565f366352

Observation 1dd6c8d1-c13f-468b-987a-f6e4568ca8da · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.896616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.071089Z digest=sha256:807e3bf118ecef42504a350317603b528310fe18edfdabb9e86d52727b829cb2

Observation 3875f7e9-9ea6-48d8-bd22-2b184af8a91f · outbound

This paper cites and Shi, B.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition and Shi, B

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.886581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.074951Z digest=sha256:425783c35880fc0b4d8de500d4d16390a2c5473cd2fe3339b2bd98b215fe313c

Observation ac8f7845-fe02-42d2-83cf-0fa56b0f684d · outbound

This paper cites H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.078580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.078580Z digest=sha256:47d10ef0f2ceb3fd7b86e2e9981ea125d93cea9d02420cef7b4174f5e7e2a3e2

Observation 6f6de3fe-b46c-4d9f-9d0c-4bcc01d00591 · outbound

This paper cites N., Zhang, Y., and Beaufays, F.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Zhang, Y., and Beaufays, F

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.869924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.082306Z digest=sha256:f392d7dececd5ef52c1c8f0964177d58f8e07a4441cbdc6895cca3f94b8769ca

Observation 976eb8dc-4005-4619-8d18-2970d5df3951 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.859640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.085973Z digest=sha256:8a19117a7a6075f95150414645d7ee32bef865286d45383c089286183d5d9a16

Observation 8bddd93f-c46e-47c0-93bc-95c37467a604 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.089684Z digest=sha256:1022eb379872dd33a8a3d0b397f6f3ed20c4228dfe32e5614b530c2b46aebed4

Observation 3cc7802d-eaff-464b-b122-a172b1f036fb · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.839686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.093433Z digest=sha256:9f9fd9f5ed37836ada5ffb186cfe3196e3134c42556212d3261976b4ecfa7ab9

Observation 83791601-00d5-48e6-b597-166a6fcf1351 · outbound

This paper cites A., Jordan, M.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A., Jordan, M

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.097176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.097176Z digest=sha256:237c566382a6b8f35c1e840ff3a59675671f865dc226f9454b1274ab20ada3bb

Observation 9c769899-7176-4dfd-a205-9ec91ae2553b · outbound

This paper cites Mixtral of Experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mixtral of Experts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.100811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.100811Z digest=sha256:21517341edcdec8a247c0c529ea7dc39ee656101c97e0701ff35a701d65560be

Observation 1b155f64-0f2f-48f0-9979-835e89202665 · outbound

This paper cites an unresolved cited work.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:46:12.824050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.104768Z digest=sha256:7bc1fc30241e98ff8c073d8efb56a338fd5adbef8bf7b963e55bfe0fb828c6b0

Observation 3c730175-c11c-4cf6-8f99-b4b2550e4145 · outbound

This paper cites Scaling Laws for Neural Language Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling Laws for Neural Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.108410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.108410Z digest=sha256:3e7d5807fab069edf14626c4205914d3ff679898d30dae65e5895a8dc2b22dbd

Observation f9ba568c-e35b-46a9-a66d-5c63c2c32c62 · outbound

This paper cites Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning video temporal dynamics with cross-modal attention for robust audio-visual speech recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.814744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.112122Z digest=sha256:150ecbaf0091d639e6abd8adb8c14ffe8bbbe5e76e5a35acbfdb5b2f18327d1d

Observation eb92f129-9c51-4e93-9443-b911ae612b24 · outbound

This paper cites Multi-task corrupted prediction for learning robust audio-visual speech representation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multi-task corrupted prediction for learning robust audio-visual speech representation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.804393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.115892Z digest=sha256:7806dbe72443cd822e1e7d1f39af2bb0d31c5862d114eaed08ecc68631d732f9

Observation e50f65bf-d38b-4a7a-95f0-414b568881b1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Adam: A Method for Stochastic Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.119735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.119735Z digest=sha256:0509d5e1ac159171ab629f523b115c1152e69d01ec6f2b21ee83cc2ae735a278

Observation 9d557243-3fe3-4ea7-a398-006108a0781c · outbound

This paper cites Moai: Mixture of all intelligence for large language and vision models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Moai: Mixture of all intelligence for large language and vision models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.794162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.123375Z digest=sha256:03c356141e8e75ea17f1bc3801a1c68c2329f3c4c18769604d2cd21c77e0167f

Observation ba99a6f1-d267-4102-9903-c77f037e8c06 · outbound

This paper cites \ GS \ hard: Scaling giant models with conditional computation and automatic sharding.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition \ GS \ hard: Scaling giant models with conditional computation and automatic sharding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.783255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.127007Z digest=sha256:3bf846b859920e87ebf69d4d9c320f94b476b5a00a407c43a785ac71faad7d9d

Observation 3ad298ad-4c4a-4eb7-b2e3-256862aa31fd · outbound

This paper cites Unified cross-modal attention: Robust audio-visual speech recognition and beyond.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Unified cross-modal attention: Robust audio-visual speech recognition and beyond

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.772919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.130698Z digest=sha256:f6435da26820e553e7dda3e273e799122cfcc6e35fd9735bcd93c85dca6d81a2

Observation 982ffbce-c590-48c5-981d-0ec2a7725ecf · outbound

This paper cites Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.762175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.134232Z digest=sha256:fc8dd463bd72c557868c29b4df9f9df6dece51a614ed94fad784f359349456a4

Observation 7ee4fd34-f493-45cf-93ec-be559e2e05e9 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.137924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.137924Z digest=sha256:f87d768dec7c2a742ce97e93cc01a2cb5c96e4f5ed43c9ce28ec11af50078a67

Observation 2e1e6b9f-0e75-4527-98d2-32b203226ecd · outbound

This paper cites Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.750980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.141733Z digest=sha256:c42e55c06d9a426433735d13cb4fb5a805f95f96ae3596a0661ae9c7c53cba5b

Observation 4d5394ea-e862-476a-9d68-7db199733b31 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.145504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.145504Z digest=sha256:0ef90be324ea6b8979a7d6c8f9b49191a65886d6db582932e7c813035826b4af

Observation 76554e89-2cc7-415b-9f11-c51a54917429 · outbound

This paper cites W., and Pantic, M.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., and Pantic, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.740438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.149611Z digest=sha256:174da275b123a278cdf0980779731dbba867b587f742fc022a91f3be6103aaa5

Observation decb22d6-5180-42cb-b35a-4c68ea23fd0b · outbound

This paper cites End-to-end audio-visual speech recognition with conformers.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition End-to-end audio-visual speech recognition with conformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.730261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.153350Z digest=sha256:4778471c8417b4201851167cf4ffdbed3f0e9275ef58f0193de29c5e72cccc90

Observation b6e30b98-4524-45fd-ba64-55921ab1de61 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.719069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.156880Z digest=sha256:318e37a145503f4137870447067d4906d5ceb7af2d751e6b5330e106b7079fc5

Observation 7a4dd371-89e0-49d0-8b75-82ee9dfb957c · outbound

This paper cites Recurrent neural network transducer for audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Recurrent neural network transducer for audio-visual speech recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.707947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.160528Z digest=sha256:c2127e274f0aa0f88bc182dbcc5ddab85a76b3a0ab615c2ddda01ed745f85d19

Observation 40f518dc-ba49-4251-969c-5dbce41495e0 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.696792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.164169Z digest=sha256:b79fd099c2a644d8cb7699a97423534cf0d30bf6eff539b67f634f64a3bb9f27

Observation 736a4236-703c-4be2-8c3b-53dfb9d93e9c · outbound

This paper cites Multimodal contrastive learning with limoe: the language-image mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Multimodal contrastive learning with limoe: the language-image mixture of experts

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.686160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.167622Z digest=sha256:02a1623760819da88456ae230bce85764c8c336b9f73cfadc8b3d0f50590a472

Observation 262ab514-50ef-47a2-abb3-201a5b8f83c8 · outbound

This paper cites G., and Ogata, T.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition G., and Ogata, T

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.675558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.171272Z digest=sha256:236ca08d1f5ada624e535e43cb7608e11620a5dfcb7b151c8e7cdb61b6be24e8

Observation a554eb4e-ae2a-4261-95fa-43ec6a0329ce · outbound

This paper cites Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.665127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.174871Z digest=sha256:cf7b9e13b8ff285de23dd082f1241826c780b41795a56c1f5c49560627cc8138

Observation 3a623aaa-0e13-431f-8564-fb3387ff6617 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Bleu: a method for automatic evaluation of machine translation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.179726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.179726Z digest=sha256:282d01bd1d6cc4386ebe72bb929e50e7c237b686795adcdbdbbb800de4b28c95

Observation 7c91ae05-3531-44c0-99b7-4a1e6ee84503 · outbound

This paper cites A call for clarity in reporting bleu scores.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition A call for clarity in reporting bleu scores

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.648936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.183249Z digest=sha256:4c21b4fd92b24f3e2f5573b996bf06429adbaa520f953310ae4af4b52321cd85

Observation faf4f2ff-e926-49b8-9ebe-e85bfd014e25 · outbound

This paper cites Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.638468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.186969Z digest=sha256:bbf8045ae415faccd50d132adbbed5c261b4a2d3c4269c10c854ff86bbedee74

Observation 1b11b7af-4db4-4e2b-92e6-e9953ef6fad7 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.190741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.190741Z digest=sha256:2299960a97c451038f51c60767d82fe38cecccc559fa0c007d1fc196a06cf73a

Observation 9e558eeb-03d6-4dde-9b12-1e0885aa2506 · outbound

This paper cites Learning from the master: Distilling cross-modal advanced knowledge for lip reading.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning from the master: Distilling cross-modal advanced knowledge for lip reading

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.622031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.194385Z digest=sha256:a118d87cf7f1cd2b7f5644cd3d509322a74b887eecf7d12b3973f9f9ee206b0f

Observation e5f9b8f4-3539-49df-ba88-92dbcfee486b · outbound

This paper cites wav2vec: Unsupervised pre-training for speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition wav2vec: Unsupervised pre-training for speech recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.611683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.197809Z digest=sha256:537181fe98b22916bbd12413d01eb2e598c1d8150997d5f662465095b2227e01

Observation e6fe91de-707b-404d-b25d-68d2f3a1f931 · outbound

This paper cites H., Nagrani, A., and Schmid, C.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition H., Nagrani, A., and Schmid, C

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.600769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.201418Z digest=sha256:93fba0fcf67e460f9f8a93b7efe38028ac4a50eda77ae5ea2e223b4656272e4e

Observation 93c99fbf-3002-4c1e-bd85-c855be4c7cd0 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.589682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.204937Z digest=sha256:16a1bb9d6c30ffc7d17ac5e19382b7b0d33c050a656071a0250a2fee8c5feefa

Observation 698d2898-de93-452b-a09a-31664d927a91 · outbound

This paper cites Scaling vision-language models with sparse mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Scaling vision-language models with sparse mixture of experts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.579332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.208750Z digest=sha256:4a51f1afd163c4b502289d60df426d3069016d1180ef80a50570cd6d42474faa

Observation 761244b3-694b-40ea-b43e-498b303e25b9 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Learning audio-visual speech representation by masked multimodal cluster prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.567388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.212277Z digest=sha256:985e3a32843ecf6623101e148d3656885a6293bc565560f50a8212e4782de8d5

Observation 6455672e-f8dc-4faa-beec-aee8ca7f209d · outbound

This paper cites Robust self-supervised audio-visual speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust self-supervised audio-visual speech recognition

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.556400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.215914Z digest=sha256:1bf49c10c7d5d7344a9ee19b4b2f12e9bbffd22413663cef554155f49ff75e63

Observation e469726c-a996-4160-b4cf-4052afa9dfe9 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition MUSAN: A Music, Speech, and Noise Corpus

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.219928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.219928Z digest=sha256:8fc71638f096d345708d5111126b7cefa4a74e41d69ae06b6b75d7156cfa75e5

Observation ea26ca4a-a117-4a82-87d1-9e3eb8a279eb · outbound

This paper cites The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.223984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.223984Z digest=sha256:dd6d25f075e8ec19227384b5bed62916cc86e894c4270718b011c334ab501cfb

Observation 70e6dd0a-b44f-498d-b5c3-5266d235f667 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition N., Kaiser, ., and Polosukhin, I

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.227915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.227915Z digest=sha256:d95f02bbe176eb6c56e783f522e6b017dcf26f23621edbc6f2509e0bb0eb6c8d

Observation 683256ca-aa77-47d4-9a5f-4f8c754714e0 · outbound

This paper cites T., and Li, H.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition T., and Li, H

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.534091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.231696Z digest=sha256:bfcfa3187750b4c773613cf9df9a6fbe195b1d60e03e12ae1b864c9fda5a2693

Observation 1e95febc-e8b0-4589-b53d-916f1e2398be · outbound

This paper cites Language-routing mixture of experts for multilingual and code-switching speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Language-routing mixture of experts for multilingual and code-switching speech recognition

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.524068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.235459Z digest=sha256:40ea58c90a77037b0e3a2ae90d17042aebc9cd388cab28c7e1fde1af3882a7db

Observation 555f1a1d-5da3-4299-9a20-b307e07dc4fe · outbound

This paper cites R., and Hayashi, T.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition R., and Hayashi, T

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.513677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.239453Z digest=sha256:ca6cecd7c1910d250c03bfd02e10111479464b58dd369920841dab537668eb27

Observation 962a535d-9a14-443b-9c67-34f771193ca1 · outbound

This paper cites Robust audiovisual speech recognition models with mixture-of-experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Robust audiovisual speech recognition models with mixture-of-experts

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.503834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.243307Z digest=sha256:61eb5594cd401d49b1343461aaef320c495e41b29acc3e6372fa5c4ea34a22d2

Observation 3584682c-30ff-4622-9c96-728984916f11 · outbound

This paper cites Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.492396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.246899Z digest=sha256:9b236567503a71238b85e8311f4dc4e4700f999d43af97023355b74c6415a3d7

Observation d8f1c0d4-9dd2-4fcb-81c8-501a463e55d6 · outbound

This paper cites Speechmoe2: Mixture-of-experts model with improved routing.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Speechmoe2: Mixture-of-experts model with improved routing

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.477707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.250437Z digest=sha256:e0f41b46071dd739b784b824e5c7437b6344dc61870d3e89c255328ec7a48d63

Observation 08b33f54-e9a5-4b79-b0ed-562a83643f94 · outbound

This paper cites Visual hallucination elevates speech recognition.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Visual hallucination elevates speech recognition

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.464533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.253960Z digest=sha256:ae415843a1f66b949e788d5531b9e93b3b7288bdfd667589027b913a18081d59

Observation 5b13026c-fa99-4c0d-970f-b9f68f4aff78 · outbound

This paper cites Self-supervised audio-visual speech representations learning by multimodal self-distillation.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Self-supervised audio-visual speech representations learning by multimodal self-distillation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.453215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.257656Z digest=sha256:935d5342c4ef5f9efe854ee6a3607b9c3aa769e275993a87a8c4603107da0853

Observation 3515f5d7-0b53-4273-bafd-96cd12610c8b · outbound

This paper cites Uni-perceiver-moe: Learning sparse generalist models with conditional moes.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Uni-perceiver-moe: Learning sparse generalist models with conditional moes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.440286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.261199Z digest=sha256:9e1c4e98d6b94aeb8650a7f3efbba552358e3185d5f710a8fa7b7aa7b9006fe0

Observation 0a2625e8-12f8-4dad-8059-86cde6accefa · outbound

This paper cites Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:46:12.425112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:46:12.264765Z digest=sha256:8ad1712edecc9bae7be4c8829d552bcaedb21eb401abda5ce1bb8a0bf1689f78

Observation b3cea8a1-ca54-4a16-b6fc-d1420e753587 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T12:46:12.269019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:46:12.269019Z digest=sha256:d6dc5c91be047e02fefa60e2e5f1d103511c257e2493d8832de1d4c3c157770f

Pith citing papers

Observation 10189419-0fe1-4ad1-9f36-c209cea88af3 · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:34.345367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:39:34.111978Z digest=sha256:d8a28786ec7695525573b7561b3bb5f3c6a89e37abcc7d1ba80dedc6d01ad30d