Pith. sign in

Paper Citation Record · LEDGER

MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19037 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:08:19.017587Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:45:51.528645Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T06:11:01.750278Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3e573cf-65f3-49a9-b1e5-55b9c13ca43f · outbound

This paper cites Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-15T18:08:19.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.766507Z digest=sha256:d37baae48f3cb1aef276765011123ce7dffafcd83411fc3cfa14f65f5a6f83ee

Observation f90e351b-e603-4ff9-9c94-b1c93912ee2a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.771390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.771390Z digest=sha256:9513f6d5ac2a45c9e53c29403746db65a7489f494326fee5800359fc7129f617

Observation 21ea53f2-f7e7-40b5-b809-13cb0a4817f0 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Next-gpt: Any-to-any multimodal llm,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.826497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.775950Z digest=sha256:11622c193b0f591126621b22dda68ac07c7f3d5168a015a009bd375c95727db1

Observation b42e5b2e-62da-443f-88ee-1b53fa616297 · outbound

This paper cites Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.780134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.780134Z digest=sha256:721e5f294e974f474e2d03c767fb8ecca2f16d8875311eb5fdead23b5059396e

Observation 3f8cdab4-09ad-4202-a4bb-e82f7aad9105 · outbound

This paper cites Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:08:19.194368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.784896Z digest=sha256:8e51fc054a788641581227dea08824098189780ee66592e393bf8230daf14cf4

Observation 1e7b3b71-36b7-492f-8ccf-44d5c78c4b62 · outbound

This paper cites Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.814059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.789458Z digest=sha256:6127e8e3b748acd80fe51814275e4a2679bc0919af59534ac6a2713c3c03e2cf

Observation a0b88f1b-85f6-4be0-b476-0c7d23eddc74 · outbound

This paper cites Speech recognition meets large language model: Benchmarking, models, and exploration,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speech recognition meets large language model: Benchmarking, models, and exploration,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.801207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.794719Z digest=sha256:5581f952def8c2259c394d92fda89384121e63ff0cdadbcd67313d2c97834cbf

Observation 678efdf9-8030-46c2-8235-4433d439abe0 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.786852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.799622Z digest=sha256:cfe40e1d2fd91c05c9bd8b929b3fd526226ce615bd0eed24f0e11f0dc830df77

Observation c7dd4b52-4a83-4b33-bfb7-190f8bcc8de3 · outbound

This paper cites A comprehensive review of multimodal large language models: Performance and challenges across different tasks,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? A comprehensive review of multimodal large language models: Performance and challenges across different tasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.774604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.803629Z digest=sha256:6bc60b7f35874154e07b6ea71670bd0831a21fe4d7b583974a0351a837336487

Observation 913df5e4-5cd2-486e-8865-3de22a8503af · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.807653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.807653Z digest=sha256:d6f71fd4c00ac1fb99fbcc7ad1094f46ce642df5ab7d43331fba314b9763bfc5

Observation d24fb121-cf20-4d97-9c81-0863c290eca7 · outbound

This paper cites MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.812315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.812315Z digest=sha256:429e0c5f8782768f29e16dcf4001a4e47aca008b13621ce1668f088cb28a406d

Observation 07b3c3f9-e185-4528-8319-63a29e8699c4 · outbound

This paper cites Large lan- guage models are strong audio-visual speech recognition learners,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are strong audio-visual speech recognition learners,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.760517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.816674Z digest=sha256:aff44215fd28e8e234c2ad20bd3e9f99add4c18759c008143ba711cdce18ea0b

Observation 0d0b914e-cf80-4843-919d-bcef5a4926cf · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.747066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.820439Z digest=sha256:d155ab00e0bee1d28f29b21190eb808d846a36dc75191de1dd3b31af5fa9a5d8

Observation d1e39b58-d6a1-443a-9657-d0a941bdac5f · outbound

This paper cites Large lan- guage models are efficient learners of noise-robust speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are efficient learners of noise-robust speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.733800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.824792Z digest=sha256:deab757ee6090c343f2be6140f97d1337a1c0d66fba1541709e5cf453866d244

Observation 07648979-b6d1-4176-bf51-4eba9e608b8a · outbound

This paper cites Avatar: Un- constrained audiovisual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avatar: Un- constrained audiovisual speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.720447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.829101Z digest=sha256:e20ff6510a35e58ac9516cb50ad30904b6c41cfe414f4c51edd255b42438ecac

Observation a779a849-e73d-4472-b0b9-d18e1f4b451c · outbound

This paper cites Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.707244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.833367Z digest=sha256:5c33842b609c477381fcb32d101065bbc542fd598552a0b1e18f295399c0587c

Observation 6cbf5381-b990-4157-ab70-8e85debd40f0 · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba: Linear-time sequence mod- eling with selective state spaces,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.694982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.837512Z digest=sha256:948460d19b451f1225ab4f4d4ff616e2bb7c113fc781b6d3d1043c36340f1c78

Observation 9454dc90-f285-4674-9527-e0c0beed982c · outbound

This paper cites Improved baselines with visual instruction tuning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Improved baselines with visual instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.841709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.841709Z digest=sha256:cbf71efdf2922d0dc4498c205ba744df5e9c684418011c4da9988a6c38af4055

Observation d51339a1-297b-4b06-8b8c-81d0a4500bdf · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.845694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.845694Z digest=sha256:ad8b0e528e94320f40a066880c77c7debdf44d86b59625b24e8c65e871aaee6b

Observation d81f9b08-7c84-4b94-a780-d72efe7c4c06 · outbound

This paper cites Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.674964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.850564Z digest=sha256:5f13d6d66b5fc831ee5a1a05a01ccb4ed9e58f89d382e21631c892f6781e5353

Observation aaf9ea8d-38c8-43ae-bf8b-257ce7e9322b · outbound

This paper cites Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.660556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.856094Z digest=sha256:b2198daa14b474439defcb115900d425afe3b5bb7c9da98ab28daf4629da075c

Observation 6587fa31-2a7e-4651-b08e-900b35ef08b0 · outbound

This paper cites Cobra: Extending mamba to multi-modal large language model for efficient inference,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cobra: Extending mamba to multi-modal large language model for efficient inference,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.648770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.860457Z digest=sha256:20fe9042048ea54185757dd5579a4105f353a86b5a6fe8d2f0f7766dcc50a0e2

Observation 10d2330a-fdf2-48ab-a2a9-e37e4f767d12 · outbound

This paper cites U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.864679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.864679Z digest=sha256:45613da4cbe9fbd983be18f80640a61138218b24659f8203659bdeec13f4ba48

Observation b7637d7e-45e8-4168-a67f-e961fdc2d588 · outbound

This paper cites Vl-mamba: Exploring state space models for multimodal learning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Vl-mamba: Exploring state space models for multimodal learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.637526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.868502Z digest=sha256:3a4bbbe6e221f73e8b0f7bea25a1b0ad2c0a3915151ed9f3731081992af2926a

Observation c1523190-95e2-410a-8d53-10aacc321d34 · outbound

This paper cites Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.625715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.872168Z digest=sha256:b5eea45d5dae9df9f79ec69e524b176c48460aa590bc61f87274475c21b2c982

Observation 3636cc29-cd3c-4e8c-94e0-57b2dabd6e74 · outbound

This paper cites Diffusion- lm improves controllable text generation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Diffusion- lm improves controllable text generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.612881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.885465Z digest=sha256:872d89728762c7cb845b092433aa20ff8e871c5de3bb0ffe9bb1511c9b749ab9

Observation 5be19410-7fcc-4241-916c-2b76a9b8247d · outbound

This paper cites Large language diffusion models,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large language diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.600420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.890194Z digest=sha256:8f6d6fafb85792eabd87f7510ba988de2d6ce577e3d041771876084cf2c54ace

Observation 272cceb3-fdc4-461b-bff8-8b12b6551c15 · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.589094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.895150Z digest=sha256:9acb4469c3182a2878cc127ee15c64e2043f46220a4830e2258f69f038f1d41e

Observation 2d55106d-d58e-4a28-934b-ad3fc2bedbe8 · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba in speech: Towards an alternative to self-attention,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.575908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.899530Z digest=sha256:ec35498ee111042be2379e4bf707b0b7403cd1b08b646f903c9cd37cabf803da

Observation fb5446a7-e0b9-4e22-90fd-ab3eb3f0b433 · outbound

This paper cites Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.564962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.903286Z digest=sha256:5d20134ae56037f6017671da03e0cf967d99c7fdaa2214ff3a91be6b991b5073

Observation fb09a6e2-85e7-488a-aa59-e4d6fb97ec9f · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.907558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.907558Z digest=sha256:5b62ba7f12b9b5204e7ed8800b7625e7d3a70e283f2318796434f60c2c91411d

Observation b50ac2db-981b-41c4-8db5-b16c887c2a86 · outbound

This paper cites Avformer: Injecting vision into frozen speech models for zero-shot av-asr,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avformer: Injecting vision into frozen speech models for zero-shot av-asr,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.553143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.911702Z digest=sha256:9b40fe04faaaacb6c839816343658c787bb153fc4ce086cda5b8743f31c85238

Observation b793b600-57fd-4929-b18c-e0f972b4440a · outbound

This paper cites Gesture-aware zero-shot speech recognition for patients with language disorders,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gesture-aware zero-shot speech recognition for patients with language disorders,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.541550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.916276Z digest=sha256:831514d0810009885fce171d63d49d1598ed7b10d70a218b6514bdac0ab362dd

Observation 2792da03-5baf-4951-a1fc-32dca1273452 · outbound

This paper cites Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.920776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.920776Z digest=sha256:7d4750c865502a6ad89903d6b457c2bd852b5111bfefb6cca92a39d4082a7347

Observation fc9ea407-1853-49bf-9f5d-8d534cac714e · outbound

This paper cites Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.924921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.924921Z digest=sha256:38b6fa8e702579bbaef0ed0791357fb8c5554c418375e48ce5abca2e69d9190d

Observation 00be9c3e-7574-41a0-bc77-9ec09d650be3 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Robust speech recognition via large-scale weak supervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.529812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.929951Z digest=sha256:092c47c4d4a661216748f05d212a741aaa972822cdb8ba81bfe39232842cd569

Observation 47ed3191-fbb7-4a33-8c57-74b2f4f57603 · outbound

This paper cites Perceptual score: What data modalities does your model perceive?.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Perceptual score: What data modalities does your model perceive?

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.518476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.933403Z digest=sha256:9a627d58b80726f99c28a96a6b7fa2ac847b8f35b098ce82515976124a2402be

Observation 68229392-882f-44df-8fb1-1c45af54fd8f · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? OPT: Open Pre-trained Transformer Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.936763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.936763Z digest=sha256:1e2c5221cc6218e16b77505eacd20b70e1b2020d7264caa32e8e68c5b2882a3d

Observation 4445125b-d4d2-4759-b990-34cf2ae851af · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Librispeech: an asr corpus based on public domain audio books,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.507533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.940735Z digest=sha256:27647da2c6bf889614b2acb077afcb5f44e4b179642dd6dbf623cf4477d3b21e

Observation 94bb2653-48d2-44cb-9284-6e7eebf93853 · outbound

This paper cites Cvss corpus and massively multilingual speech-to- speech translation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cvss corpus and massively multilingual speech-to- speech translation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.496644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.945358Z digest=sha256:2ffbe6db6d923eb805518d7dac1f2fbd8184f9f9d685d74e13ac9b09a28a57f7

Observation 6654fe36-8e69-4982-97bd-87b4c9347b63 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.949250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.949250Z digest=sha256:bfc00585c76e00134a69038d3387f1a56f4aa9ec86a9ce39c5014747325b25a0

Observation 408f3226-4033-4374-8f2e-e8539d49e997 · outbound

This paper cites Microsoft coco: Common objects in context,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Microsoft coco: Common objects in context,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.485293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.953240Z digest=sha256:a35823e443d2e33c5d5fe94ffcc2336701b025d0c20741b34f8d3f9c3553e5de

Observation c239201b-f9af-417f-8de9-6f4f4bb0558e · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.957514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.957514Z digest=sha256:76cebfd28c9bfe6891837a3ea09745846b6321cac406c3c7bf6988565ea82720

Observation f61914bb-bbae-4dcc-8c67-944074878096 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.472960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.961349Z digest=sha256:568db03c5154dee1b9ce10578f8ef1352c9e62b601b93518fb3f9145aaf8d678

Observation d0de09d2-5e37-47ea-a87c-4aafa2b5146a · outbound

This paper cites Zero-shot text- to-image generation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-shot text- to-image generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.461940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.965387Z digest=sha256:0f9bbb533c7d9ed63b172cb6375f52a51ecdc6f0a3ae464c2288de4214315312

Observation e84adce2-67ee-4889-b71d-f08a113f675e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.450637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.969045Z digest=sha256:b2ca2d0c7416f5047c356a3198801431ae25fdd401804b8c701a57df1c9a4dce

Observation 34763197-be4c-459c-9712-e0ff39790d65 · outbound

This paper cites Decoupled weight decay regularization,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Decoupled weight decay regularization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.437858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.972445Z digest=sha256:5ab044e599849114195228e18d81c14583b2caa7ec80b940817a4a36a364fba3

Observation 9ad27258-a6d2-41f0-960b-2870186be976 · outbound

This paper cites Deep audio- visual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Deep audio- visual speech recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.426624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.976171Z digest=sha256:748e5794dd83c85d7cc34e5b6d61cf6003979eb9418e167369afa1d6b14a40b9

Observation 4774422d-0fbb-4c12-87da-50305f889e5a · outbound

This paper cites Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.414781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.980618Z digest=sha256:05c3ef1f9d8b58253994aed4c68675269be7ebc2e8cd9bf963cb115af9893140

Observation 4f1d409a-8add-4c12-93ca-11ffb967f691 · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Spoken moments: Learning joint audio-visual representations from video descriptions,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.402749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.984832Z digest=sha256:30284454993610f962b97977c3214df6a10b10735f75895feedd2d1881f2195f

Observation 8f5bff19-e950-4cb5-a898-7b14a40d218b · outbound

This paper cites pyttsx3: Offline text to speech (tts) converter for python,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? pyttsx3: Offline text to speech (tts) converter for python,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.390891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.988723Z digest=sha256:432a57f04c915f1aed21e94eb3315c16261a9fa3e3907885e5f9014317652261

Observation 502a7683-4b4f-4e05-b863-413ee6eca90a · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MUSAN: A Music, Speech, and Noise Corpus

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.992534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.992534Z digest=sha256:2c4a2200554d99b98a3acc28eb4917cc17102a6ea340f68970f45feeda528877

Observation 89498a63-f6cb-43e0-b679-fece7763bde2 · outbound

This paper cites Easyocr: Ready-to-use ocr with 80+ supported lan- guages,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Easyocr: Ready-to-use ocr with 80+ supported lan- guages,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.377923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:18.996449Z digest=sha256:5d2c766a1bdd2c4f728aa9ef08d8561f6086303ef2c6ba87f28e0c82bfea82a0

Observation 93f3b218-634d-478c-9940-26d54c6e24bd · outbound

This paper cites Multi- moments in time: Learning and interpreting models for multi-action video understanding,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- moments in time: Learning and interpreting models for multi-action video understanding,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.366134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:19.000266Z digest=sha256:d1e1c26291fc71bf5bb1e107f6b850e7885f88359a0c546b482a1084152e89fe

Observation cc02efc9-5ce5-420c-9be2-a3f62988bd0a · outbound

This paper cites Auto- avsr: Audio-visual speech recognition with automatic labels,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Auto- avsr: Audio-visual speech recognition with automatic labels,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.353708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:19.005817Z digest=sha256:7f42d7054adc48637e97cd5912bd1dd3f53511ad2139b99162a1aac16ccbf854

Observation 8c26f861-47c6-4525-9ff7-878f0c730255 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Flamingo: a visual language model for few-shot learning,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.341067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:19.009408Z digest=sha256:bb3cc84908c03c293857d5ab470bb43a128df7b5d5cfedee225256480df75368

Observation 89a83efc-6b20-4894-a2fa-2bca2d5eb5c8 · outbound

This paper cites Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:19.013669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:19.013669Z digest=sha256:7657fd19a706358b890ecf04923396ef24f08cca9f7908cf9535c7a1c22b33ff

Observation 6ead41f1-829e-4bb4-bb49-02985ed27086 · outbound

This paper cites Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.329115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:08:19.017587Z digest=sha256:3770dd09f81f9accbcf65cdee2193a4d692e8972440f26439b952a712ba48508

Pith citing papers

Observation 19493417-f083-4e90-ac6b-3d3122752e56 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-10T01:09:09.297559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:466795013f270641b28868052ea09c9ada5347b244a27cc398f1557fe8e4cb72