Pith. sign in

Paper Citation Record · LEDGER

MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19037 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:08:19.017587Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:45:51.528645Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T06:11:01.750278Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3e573cf-65f3-49a9-b1e5-55b9c13ca43f · outbound

This paper cites Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-15T18:08:19.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.766507Z digest=sha256:a556e4d37be3c2b475822ad2c4f3594787d73f2e74c993d9e3df7cbda74d0bf7

Observation f90e351b-e603-4ff9-9c94-b1c93912ee2a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.771390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.771390Z digest=sha256:d32e4db7017f93a6eef3d5b941b86c411dba59d7b130f39a71cbde15a3a536c4

Observation 21ea53f2-f7e7-40b5-b809-13cb0a4817f0 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Next-gpt: Any-to-any multimodal llm,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.826497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.775950Z digest=sha256:b6e6b5be2f9f8ab2d124267e00f216d20aac35b50205ac52f66e3617470477c7

Observation b42e5b2e-62da-443f-88ee-1b53fa616297 · outbound

This paper cites Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.780134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.780134Z digest=sha256:5ca59cd736fde50aff0801c7464cb72d27c9c957c7ecf5f276361ffae7449832

Observation 3f8cdab4-09ad-4202-a4bb-e82f7aad9105 · outbound

This paper cites Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:08:19.194368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.784896Z digest=sha256:f291f892b6414639d5e90b28d94315f85d8db6cee39ddf4700729251c8590ed9

Observation 1e7b3b71-36b7-492f-8ccf-44d5c78c4b62 · outbound

This paper cites Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.814059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.789458Z digest=sha256:e06d5a05bcc0c5612e44e26698238deda39ecda7ff1989c2b48b49bb1379b07d

Observation a0b88f1b-85f6-4be0-b476-0c7d23eddc74 · outbound

This paper cites Speech recognition meets large language model: Benchmarking, models, and exploration,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speech recognition meets large language model: Benchmarking, models, and exploration,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.801207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.794719Z digest=sha256:9b9dfc7d751ed035182a3f54a599cc59e56f57de39233d90769216437cca97e5

Observation 678efdf9-8030-46c2-8235-4433d439abe0 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.786852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.799622Z digest=sha256:f3555dc753062c593a49e4aa0e5721e9d6784fe1418cd363da89535954804c75

Observation c7dd4b52-4a83-4b33-bfb7-190f8bcc8de3 · outbound

This paper cites A comprehensive review of multimodal large language models: Performance and challenges across different tasks,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? A comprehensive review of multimodal large language models: Performance and challenges across different tasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.774604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.803629Z digest=sha256:6d2d5a0d31f4c7cb1b2ada63ef248e2d123fa3c7422547dae156ec0a39ba3fc0

Observation 913df5e4-5cd2-486e-8865-3de22a8503af · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.807653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.807653Z digest=sha256:28d94d5eb1a7f940fcfa9c67192a0fc136d838d3b22292b894e9459334e7a4c7

Observation d24fb121-cf20-4d97-9c81-0863c290eca7 · outbound

This paper cites MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.812315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.812315Z digest=sha256:a03bf65abc3d1581a2cadb848115af7edf61a180bcc980b7ee8e2498d813cf71

Observation 07b3c3f9-e185-4528-8319-63a29e8699c4 · outbound

This paper cites Large lan- guage models are strong audio-visual speech recognition learners,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are strong audio-visual speech recognition learners,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.760517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.816674Z digest=sha256:b1a56c2b363e9243cc9f8ad12c5b3eb5b43100af4c1f2edadc1ee7614c07a0ca

Observation 0d0b914e-cf80-4843-919d-bcef5a4926cf · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.747066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.820439Z digest=sha256:cd3ad0ad54d4161bc15a97e2d0a00cbad42e04ced6bc976675d01d491eb400ad

Observation d1e39b58-d6a1-443a-9657-d0a941bdac5f · outbound

This paper cites Large lan- guage models are efficient learners of noise-robust speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are efficient learners of noise-robust speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.733800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.824792Z digest=sha256:d9583306031698b82a9da29a8dd2b1dfd7bba087984788c5747c1c05a96d3541

Observation 07648979-b6d1-4176-bf51-4eba9e608b8a · outbound

This paper cites Avatar: Un- constrained audiovisual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avatar: Un- constrained audiovisual speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.720447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.829101Z digest=sha256:48e679886a82e625f749df4c5e7116425c4befef84d64ea942b8cc989838967d

Observation a779a849-e73d-4472-b0b9-d18e1f4b451c · outbound

This paper cites Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.707244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.833367Z digest=sha256:696876b3ed207170f2bfa1972ce2b9c0c40192c314756c20d20e179a39304fe8

Observation 6cbf5381-b990-4157-ab70-8e85debd40f0 · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba: Linear-time sequence mod- eling with selective state spaces,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.694982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.837512Z digest=sha256:e06361cd8ba504372f88eb68d190f5b674f4c224dcf11c8d892c684797a19792

Observation 9454dc90-f285-4674-9527-e0c0beed982c · outbound

This paper cites Improved baselines with visual instruction tuning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Improved baselines with visual instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.841709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.841709Z digest=sha256:cc0577b349ece5a589026deed675197f8394ec761d7bef6b0c29a23fed981ae2

Observation d51339a1-297b-4b06-8b8c-81d0a4500bdf · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.845694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.845694Z digest=sha256:bb1c40200a45fd70b3171e9d6a18352cabf3bb60801bb6fd5b407f1170a58e84

Observation d81f9b08-7c84-4b94-a780-d72efe7c4c06 · outbound

This paper cites Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.674964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.850564Z digest=sha256:69fa60192dcd2a5e54b35541d94b49c31579d043cea1b7ab2e70bc7ae3a31c98

Observation aaf9ea8d-38c8-43ae-bf8b-257ce7e9322b · outbound

This paper cites Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.660556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.856094Z digest=sha256:87dc7e1fd633aaf71d93c256589856b1411eca06b9f0176aec3cf22f29a515c0

Observation 6587fa31-2a7e-4651-b08e-900b35ef08b0 · outbound

This paper cites Cobra: Extending mamba to multi-modal large language model for efficient inference,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cobra: Extending mamba to multi-modal large language model for efficient inference,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.648770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.860457Z digest=sha256:94e168cb8679a6c6420cc16515fe923ac004687cf8161230cf4ef70629ab2fad

Observation 10d2330a-fdf2-48ab-a2a9-e37e4f767d12 · outbound

This paper cites U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.864679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.864679Z digest=sha256:0dc33c7f414c69ae87e0a7c6d15068c7e5dc60ddd510e1cc8fdb4b5a8c600769

Observation b7637d7e-45e8-4168-a67f-e961fdc2d588 · outbound

This paper cites Vl-mamba: Exploring state space models for multimodal learning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Vl-mamba: Exploring state space models for multimodal learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.637526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.868502Z digest=sha256:a32277d9c353c330ab664a3a7e2055f0ab46f7313e9bd4b92cc4bc03f3b7d9a3

Observation c1523190-95e2-410a-8d53-10aacc321d34 · outbound

This paper cites Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.625715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.872168Z digest=sha256:27c7c82da38033bde2030717791d727c572b77ccba5c650d3ccac92c753f8313

Observation 3636cc29-cd3c-4e8c-94e0-57b2dabd6e74 · outbound

This paper cites Diffusion- lm improves controllable text generation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Diffusion- lm improves controllable text generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.612881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.885465Z digest=sha256:8a9719c1669714e1d6d4ff50924abac96c8e9a828a2c131088789bdf186452ba

Observation 5be19410-7fcc-4241-916c-2b76a9b8247d · outbound

This paper cites Large language diffusion models,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large language diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.600420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.890194Z digest=sha256:f73a733be04197f38fe379ba24e19eb0b51291c8bbec0fc768320d698d2f7538

Observation 272cceb3-fdc4-461b-bff8-8b12b6551c15 · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.589094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.895150Z digest=sha256:e6cd3630645084301fc3deaea39b86dc6cfd8954767b9acf6484e5b21a1b7676

Observation 2d55106d-d58e-4a28-934b-ad3fc2bedbe8 · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba in speech: Towards an alternative to self-attention,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.575908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.899530Z digest=sha256:84a1bbfdc2d9c2f60ab7a9f613f4c2d0e236b4a8312855201a4b6aa4512f9652

Observation fb5446a7-e0b9-4e22-90fd-ab3eb3f0b433 · outbound

This paper cites Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.564962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.903286Z digest=sha256:274707afc1332e331e8d652c7248fcf615e2e3c02f4441f7f36f2e84a89bde62

Observation fb09a6e2-85e7-488a-aa59-e4d6fb97ec9f · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.907558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.907558Z digest=sha256:e7104009d195704274fff49e43d67285b35707d922c7db7a053a6e40e892d5da

Observation b50ac2db-981b-41c4-8db5-b16c887c2a86 · outbound

This paper cites Avformer: Injecting vision into frozen speech models for zero-shot av-asr,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avformer: Injecting vision into frozen speech models for zero-shot av-asr,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.553143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.911702Z digest=sha256:eaec78a1e13ca8850e54f647a1da59eede5c73e9092ed8d44a6600e0419b3e82

Observation b793b600-57fd-4929-b18c-e0f972b4440a · outbound

This paper cites Gesture-aware zero-shot speech recognition for patients with language disorders,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gesture-aware zero-shot speech recognition for patients with language disorders,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.541550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.916276Z digest=sha256:6c8ff393f34ed8250a2518d4b8a995e6eac7cab99d5a5ecaa2f611e7d27fd6d7

Observation 2792da03-5baf-4951-a1fc-32dca1273452 · outbound

This paper cites Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.920776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.920776Z digest=sha256:d13e7325b4f9d3a395f09b62c7434549509afe28c83269808c8f06f9ec24a638

Observation fc9ea407-1853-49bf-9f5d-8d534cac714e · outbound

This paper cites Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.924921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.924921Z digest=sha256:afc8408eae5a8427ad4230d1c55288c25e108c30518831d1fe9f74069af7b777

Observation 00be9c3e-7574-41a0-bc77-9ec09d650be3 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Robust speech recognition via large-scale weak supervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.529812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.929951Z digest=sha256:340c56ecbd110530fdbd82e23133d61a152abeee8f2a1f441e89953b51dd888e

Observation 47ed3191-fbb7-4a33-8c57-74b2f4f57603 · outbound

This paper cites Perceptual score: What data modalities does your model perceive?.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Perceptual score: What data modalities does your model perceive?

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.518476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.933403Z digest=sha256:ab25e0e1b14b718dc4b44740bfc134b6d7ec19a7ff06fb27d03a58515c5755fd

Observation 68229392-882f-44df-8fb1-1c45af54fd8f · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? OPT: Open Pre-trained Transformer Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.936763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.936763Z digest=sha256:42cd532f32444952dd2f429f72206f49cf6b85d7fa9ca099d3a2a8b0fc4997b7

Observation 4445125b-d4d2-4759-b990-34cf2ae851af · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Librispeech: an asr corpus based on public domain audio books,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.507533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.940735Z digest=sha256:5bab074053aae0db5923baf68a1f4955849936076d6fbc06c934f0f3b5e0ec6b

Observation 94bb2653-48d2-44cb-9284-6e7eebf93853 · outbound

This paper cites Cvss corpus and massively multilingual speech-to- speech translation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cvss corpus and massively multilingual speech-to- speech translation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.496644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.945358Z digest=sha256:a5fee3c0a8ef2782d8416f84d747413ff9084443f1c4578b82e7ca81242f2872

Observation 6654fe36-8e69-4982-97bd-87b4c9347b63 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.949250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.949250Z digest=sha256:1571e59ec2dc9b33f47fbfe52abaf291292d1a1b1d45ad274cb063c89d70795f

Observation 408f3226-4033-4374-8f2e-e8539d49e997 · outbound

This paper cites Microsoft coco: Common objects in context,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Microsoft coco: Common objects in context,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.485293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.953240Z digest=sha256:7c98d03ae4321b61f19d21e36492d38d2ec294ac441baa9163a7f8ad57cc55ed

Observation c239201b-f9af-417f-8de9-6f4f4bb0558e · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.957514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.957514Z digest=sha256:acfad0c547df7f5b932862a3b535b003314036d867b15966ee3fca856c302472

Observation f61914bb-bbae-4dcc-8c67-944074878096 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.472960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.961349Z digest=sha256:2539d5595c80f32ba006a1f0805812dfaaf6df17c6ac4b418a61c17915a1f50e

Observation d0de09d2-5e37-47ea-a87c-4aafa2b5146a · outbound

This paper cites Zero-shot text- to-image generation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-shot text- to-image generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.461940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.965387Z digest=sha256:3dfb2fa065c8c3b87934b390d4a1ba30e7a234dfe26a510b744115c72be7fb90

Observation e84adce2-67ee-4889-b71d-f08a113f675e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.450637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.969045Z digest=sha256:73e73461fbba2838dbcd68403dc7f2e7d870431688e9ddbc2edda5b44afa87ba

Observation 34763197-be4c-459c-9712-e0ff39790d65 · outbound

This paper cites Decoupled weight decay regularization,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Decoupled weight decay regularization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.437858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.972445Z digest=sha256:fcf023e5c3b416c22bc4c51830cdefa790665c97c411fe0ab0c774ff3c6b9cb8

Observation 9ad27258-a6d2-41f0-960b-2870186be976 · outbound

This paper cites Deep audio- visual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Deep audio- visual speech recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.426624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.976171Z digest=sha256:5cd7de39a29ca3329f88d1a44bdf0a285c7436e5dcb8986e947b98169dd34f0b

Observation 4774422d-0fbb-4c12-87da-50305f889e5a · outbound

This paper cites Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.414781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.980618Z digest=sha256:79bbda78b0540cc0961e4af0779d25c0f584e49b2ede0f498533d9e1d668c55c

Observation 4f1d409a-8add-4c12-93ca-11ffb967f691 · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Spoken moments: Learning joint audio-visual representations from video descriptions,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.402749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.984832Z digest=sha256:7c7789c6a047ccafde5f061def0df4266c9e63b0efca68eaeded3d7fb29b9b58

Observation 8f5bff19-e950-4cb5-a898-7b14a40d218b · outbound

This paper cites pyttsx3: Offline text to speech (tts) converter for python,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? pyttsx3: Offline text to speech (tts) converter for python,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.390891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.988723Z digest=sha256:b8dfac7b53a82e84e1680f6ef0c94ba1105f38de5f921cba97f58ce43e3c7427

Observation 502a7683-4b4f-4e05-b863-413ee6eca90a · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MUSAN: A Music, Speech, and Noise Corpus

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.992534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.992534Z digest=sha256:d3da37c88b1fc4d7ec455523d4bd622e0d740ce9a5740248d085b3cadcc21dd8

Observation 89498a63-f6cb-43e0-b679-fece7763bde2 · outbound

This paper cites Easyocr: Ready-to-use ocr with 80+ supported lan- guages,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Easyocr: Ready-to-use ocr with 80+ supported lan- guages,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.377923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:18.996449Z digest=sha256:e3ed6e8491d5b1d2729f54dea81f49d577aea121a8db077cb79a1e7d73be32f4

Observation 93f3b218-634d-478c-9940-26d54c6e24bd · outbound

This paper cites Multi- moments in time: Learning and interpreting models for multi-action video understanding,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- moments in time: Learning and interpreting models for multi-action video understanding,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.366134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:19.000266Z digest=sha256:e1758ff9abe3236170634bea182d1f1e2c8eac6babb0f9cfdac6826b89c1c985

Observation cc02efc9-5ce5-420c-9be2-a3f62988bd0a · outbound

This paper cites Auto- avsr: Audio-visual speech recognition with automatic labels,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Auto- avsr: Audio-visual speech recognition with automatic labels,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.353708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:19.005817Z digest=sha256:b6c56f2d809f9d2b3b6d351f5a2a6810b89c39fa466938b7437921c29cfc66d4

Observation 8c26f861-47c6-4525-9ff7-878f0c730255 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Flamingo: a visual language model for few-shot learning,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.341067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:19.009408Z digest=sha256:627df4f57cf97fc9b42db45ac3ec0600e860e83d001a4bb24e33b4115ec31236

Observation 89a83efc-6b20-4894-a2fa-2bca2d5eb5c8 · outbound

This paper cites Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:19.013669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:19.013669Z digest=sha256:90c88d6e93350d652457320f4260d04a458df1cfcb07032416b7bab301f36713

Observation 6ead41f1-829e-4bb4-bb49-02985ed27086 · outbound

This paper cites Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.329115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:08:19.017587Z digest=sha256:ebd343e83770ad4a59bb2eb3b6e4b9a2f0f91c8ff04fd81f1bc8734b8cbdea6e

Pith citing papers

Observation 19493417-f083-4e90-ac6b-3d3122752e56 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-10T01:09:09.297559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:5e9861aad6cf46567a53862cfcacaa4ebe3ad59629d55c54f8778d10d0d1a911