Pith. sign in

Paper Citation Record · LEDGER

mSLAM: Massively multilingual joint pre-training for speech and text

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2202.01374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.01374 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:49.835256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T21:51:17.752402Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a27e7da-de70-4ff1-aa85-a9c294b12dcf · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language mSLAM: Massively multilingual joint pre-training for speech and text

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.775223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:dfd0051969d7eea742cffd073dc37180008267f0a39219450ca2ddbaa1797190

Observation 46595cb1-d75f-46f7-8141-624fc939a5c5 · inbound

AudioPaLM: A Large Language Model That Can Speak and Listen cites this paper.

AudioPaLM: A Large Language Model That Can Speak and Listen mSLAM: Massively multilingual joint pre-training for speech and text

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:57.851683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T07:07:57.800866Z digest=sha256:3c687d56a6707bff2135e52e278ef55a1f9e8f4caabc50a18d403d330e3d4e5a

Observation beb416f4-df7a-44e0-8960-ad83b6756b07 · inbound

STORM: Strategic Orchestration of Modalities for Rare Event Classification cites this paper.

STORM: Strategic Orchestration of Modalities for Rare Event Classification mSLAM: Massively multilingual joint pre-training for speech and text

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:09:14.792158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:09:14.792158Z digest=sha256:05732e605e2a434484920bbf1807e59de74e675443adbf547d3b803b736c906f

Observation ee707da0-40d7-407b-a982-46099aa5bfa7 · inbound

Large Concept Models: Language Modeling in a Sentence Representation Space cites this paper.

Large Concept Models: Language Modeling in a Sentence Representation Space mSLAM: Massively multilingual joint pre-training for speech and text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:01.021194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:37:01.021194Z digest=sha256:622ec90de1b8f37e506570d51c9ce4ee3be8ee919dc1c2d1fadfd1f1205cda13

Observation 3ffc3bae-0de9-4ff4-8b5e-d835bb75984d · inbound

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning cites this paper.

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning mSLAM: Massively multilingual joint pre-training for speech and text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:22.389557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:27:22.389557Z digest=sha256:0db3bc3075b6196250f82a5d66aeaa9b2e67792b0f14ff6f5f2c82d229b6e141

Observation d78d5ae2-659d-4f9f-ac33-a6669d093d25 · inbound

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators cites this paper.

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators mSLAM: Massively multilingual joint pre-training for speech and text

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T12:30:52.016897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:30:52.016897Z digest=sha256:53deb5800edb822a6972f65ebf86b9a924f854420e1010eab3d49182852645bd

Observation 073b2990-7cce-476c-be3a-76ec126d729b · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment mSLAM: Massively multilingual joint pre-training for speech and text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.929162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.929162Z digest=sha256:d09a532eb9f7aa3fbe95e9022679c889a6f88792e8d47d9e33bd0b59a9647841

Observation 21895f18-f986-4ed7-aacf-4a66dc443483 · inbound

Self-Improvement for Audio Large Language Model using Unlabeled Speech cites this paper.

Self-Improvement for Audio Large Language Model using Unlabeled Speech mSLAM: Massively multilingual joint pre-training for speech and text

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:49.835256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:49.835256Z digest=sha256:f7b6010821a1b38bfd8eede686ceaac6387a485a1b521bd0f589588dc3f0e47a

Observation d12b6ab0-06fa-4b7e-8094-9df5092d3073 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs mSLAM: Massively multilingual joint pre-training for speech and text

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.754099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:354b939f90a32de733c278b8061c2b1ddc9fead95d76905179e0024fd3b54ba5

Observation 78517cd0-5e7d-41bc-be3c-0882b3a65c71 · inbound

Text-Utilization for Encoder-dominated Speech Recognition Models cites this paper.

Text-Utilization for Encoder-dominated Speech Recognition Models mSLAM: Massively multilingual joint pre-training for speech and text

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:51:25.132172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T13:37:09.614784Z digest=sha256:81f26187245d07c63d22ed7c52941ee01f0101f1b383359769948412f7fa6b04