Pith. sign in

Paper Citation Record · LEDGER

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2412.01145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01145 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:25.985859Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:52:53.386966Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T20:45:08.147897Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 985a90aa-debf-49ca-be33-294fffdccfa5 · outbound

This paper cites Language models are few-shot learners,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Language models are few-shot learners,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.697786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.784371Z digest=sha256:847960c61087142143b12cde96c50a5d4353ef39e05e077a4ce8fa788d040efd

Observation 716d7f8d-3a7c-4636-b8df-37eb5a2e534d · outbound

This paper cites GPT-4 Technical Report.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.790256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.790256Z digest=sha256:bb83ccadc03612ba699b3cea6ceb80a955892ef2f2e9fdc7b64de250d0c193a0

Observation 0cb7d0de-fd21-46ad-abff-27c6cbd21103 · outbound

This paper cites The Llama 3 Herd of Models.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.795189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.795189Z digest=sha256:d1a01f5ba73f58b4aa431d0fd94c8285356d2650767c1814f7119fc98c3fda0b

Observation c8df4194-3025-4b92-abd6-a0202eeb08ec · outbound

This paper cites Self-instruct: Aligning language models with self- generated instructions,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Self-instruct: Aligning language models with self- generated instructions,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.680275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.801322Z digest=sha256:8b45dbcec7a61d2bf6017c317384b8d9cf571d9c2ea1f2797bfab68ab79d081f

Observation 3cdce924-0448-4e9e-bce3-0b7308d2b00e · outbound

This paper cites Training language models to follow instructions with human feedback,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Training language models to follow instructions with human feedback,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.806894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.806894Z digest=sha256:2dc850a7730e6f7a67da44369e091bc0e950cb8e91bb0b6f25f9eac73b8b106e

Observation 329c671c-df29-4ec5-92ad-516a771e5422 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Direct preference optimization: Your language model is secretly a reward model,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.811299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.811299Z digest=sha256:0df39d61e1cfae73cf3a41d5aa1330b1770491cf4e780893fc1ea2cb2dd17fdc

Observation 3a2cc399-9573-4da9-aefa-2645b242341f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Moshi: a speech-text foundation model for real-time dialogue

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.816937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.816937Z digest=sha256:3f3241b770ee140b5d8a187e314eca8361706fbd71557b6d2fb3c053f3fe7303

Observation 06b2145c-2d85-4191-bde8-1c5a2f153906 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.821357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.821357Z digest=sha256:fe0b6e3f1bb3bebec92eb0dcab4cdb5c9ed9dea8ec6619ea996917604071464c

Observation 36998284-953c-477a-95da-df47762a1d60 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.825642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.825642Z digest=sha256:481a0dfa37377477d71fad27b77268c785fbbd80f11f90ad99b78b0eec4d5845

Observation 8c733e4d-e0fc-41df-9404-1b8e19304a9a · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.829811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.829811Z digest=sha256:034c3892fd19acd5bd6ef82236959cf870c2a1029ef2c84e50de3fa1643c0e54

Observation 19c436fa-ff2f-4d68-8e51-6a9e950764e1 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.834345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.834345Z digest=sha256:58d5664878fcbb72743b71e152098988e65aa583438352ae27466f28062a441e

Observation fedfc8f9-0c10-4bc1-80ab-eae0ffdfc593 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.838344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.838344Z digest=sha256:aacd689545559729de8d5276668e8c69f5b78e9a9a71bef7583effc4d0f7cd89

Observation e60d1685-300a-4a71-b8d3-9d17e046fbbc · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.841973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.841973Z digest=sha256:c81742583c69a786e56c0be83f7d2fe2fbdf3dd8d29229eece3150f1e19f6cea

Observation c4f4b24e-79ee-4853-945e-ef4672dfdf7d · outbound

This paper cites SALMONN: towards generic hearing abilities for large language models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SALMONN: towards generic hearing abilities for large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.641519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.845827Z digest=sha256:f4c4c986ef4dca4e633812fca0d89fa2a5e9dcbb8e9db265960922fe4f2f6c99

Observation 5997242c-74a7-4943-81cf-1c028e5d6d06 · outbound

This paper cites Qwen2-Audio Technical Report.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Qwen2-Audio Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.849796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.849796Z digest=sha256:f1d4e7303aabd69fe6f2e040ae1e3b40b74246660f8eb340a9b91f07c172e863

Observation 181cedd0-34f3-49bf-a163-cc409242e82e · outbound

This paper cites On decoder-only architecture for speech- to-text and large language model integration,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM On decoder-only architecture for speech- to-text and large language model integration,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.626933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.854425Z digest=sha256:88f830de80fd0848a0c071abed91e622b028d7f3f5d41541ad1080978952070e

Observation 782fcbb5-d293-44d4-b352-6e7db367d78f · outbound

This paper cites COSMIC: data efficient instruction-tuning for speech in-context learn- ing,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM COSMIC: data efficient instruction-tuning for speech in-context learn- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.610104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.858028Z digest=sha256:37a2fc67217ec2542a3d24459497d4abd4145796553885a6ef5c70bb9181367b

Observation e622f2e9-4fe5-42dc-bd14-35fab6d28c79 · outbound

This paper cites Prompting large language models with speech recognition abilities,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Prompting large language models with speech recognition abilities,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.590364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.862163Z digest=sha256:5a0939b94e3f0863213289fd28c18a6f698820046b0a7cd035f5f0e85e30260b

Observation 7df69f19-523c-44be-8936-bd2a9a07ecf9 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.866877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.866877Z digest=sha256:71a96a9cfa6794a0fb9fe2f0501ce5269606a5a161aa2757320b619cd8139fc6

Observation 97d59d3c-1b84-4e41-85d6-add42654e0e4 · outbound

This paper cites High fidelity neural audio compression,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM High fidelity neural audio compression,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.870796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.870796Z digest=sha256:de96d230867b2b3072c4b6d25594c3563d2657cdaf78b9d8f3ad231f3718fd91

Observation 623da227-e956-4251-886e-484a7e084a54 · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM High- fidelity audio compression with improved rvqgan,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.563623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.874350Z digest=sha256:b0bf187d540c811a8f44ce67c77f1f5ca1d745a4ed8c8a092bbb809477f8bf96

Observation 818a281e-6dc3-40a4-a1bd-5befccdef7f2 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.878311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.878311Z digest=sha256:d1d8555a8c6e71eaa04df4b7b0586852df41323878ea25313482038842e3e10e

Observation 873c4ec5-451a-4960-a1e5-be8b46640b66 · outbound

This paper cites Wavllm: Towards robust and adaptive speech large language model,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Wavllm: Towards robust and adaptive speech large language model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.553469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.882930Z digest=sha256:69807817fefc0d0c20315b9fcb15a21bb24ec1d1c22334df521f66482aaf4cce

Observation da070468-5174-4080-9257-9d6e12c61ff6 · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Audiochatllama: Towards general-purpose speech abilities for llms,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.519938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.889898Z digest=sha256:11a84aa2d377dbe12061902bd965ff5782db8c40702079526cf1d305b620ce6d

Observation b6e663c6-107c-432c-9304-37c183a30bc2 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.893673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.893673Z digest=sha256:bcb02695c82d8bf6ed6f3d3c646b1c1c96d492a536baecf2cb581963bb4c27c2

Observation 44f31491-4a10-408e-9d3f-60af5e18d21c · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.897763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.897763Z digest=sha256:da1279a4aa04bad6240f82b9d734875e72a42aebe2c97439e854cb35e50f06d6

Observation 36749f60-17f5-410b-810f-9da282938ede · outbound

This paper cites Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.901872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.901872Z digest=sha256:0eceeb4b1119e1f6d1f1847ceb1ebe5ea89f3360504911f7bcb153b6f0ac0ccc

Observation 6c5b8b81-e1a0-439a-b105-2433776bea43 · outbound

This paper cites Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.506587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.905887Z digest=sha256:4214e9adbd0680ef9d0eaf9b1a3b8f055f58cab97796f9027c41c273f87f367f

Observation 92414057-f770-43da-ad24-19aa9404c3a0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.909775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.909775Z digest=sha256:9446dee0a43a4be5379772b47a359999b76e9cd85885578275df69919dc7dc4f

Observation 36333554-2450-44f8-8090-6b818c617999 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.914477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.914477Z digest=sha256:b9bc3c7dec410c994b129c1461fe4e9c03f70f1c4db82da03d003d3733d8c019

Observation 65626896-ee5e-4025-a331-63ae60dc2ac7 · outbound

This paper cites Connec- tionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Connec- tionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.490200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.919090Z digest=sha256:8081397ee1146da1c761961787c00679010f6af9cdfd4ca7c44291bc43627056

Observation 71511e59-350c-4d5c-a4fd-e8c8a0a96a34 · outbound

This paper cites CASS-NAT: CTC alignment- based single step non-autoregressive transformer for speech recognition,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CASS-NAT: CTC alignment- based single step non-autoregressive transformer for speech recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.473692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.923067Z digest=sha256:ee1c621ad8d2235fd86fd0b80bc7d43cf87786123e51bbd3d1eed942079c10e8

Observation e5e6650f-2ee5-4b1c-ab49-d42a5d3749a0 · outbound

This paper cites Unienc-cassnat: An encoder-only non-autoregressive asr for speech ssl models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Unienc-cassnat: An encoder-only non-autoregressive asr for speech ssl models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.461458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.927088Z digest=sha256:e221c10f909d430f25c313c6b8690f49936b10930d6ecb4c1346326a9b8e9b19

Observation 9c2802fb-909d-4e2d-b02f-3ac59ab22d71 · outbound

This paper cites Ctc-based compression for direct speech translation,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Ctc-based compression for direct speech translation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.449574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.932204Z digest=sha256:8e52d7dda521a7ee92db41bed80148fc79b12dce447b8838916a615c90b6a3fb

Observation 982d628b-f3e8-4a9b-af47-f524d73fcd9d · outbound

This paper cites CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.435311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.935963Z digest=sha256:7179f7fc3973ac7331561c75f759032ec258465fbdde073484c78eec1756a9e5

Observation a2d75a61-cb1c-437a-bc2b-b126ea13604c · outbound

This paper cites SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.421186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.939645Z digest=sha256:07173b5af69191b50da67a607ebf89fb4aa1bda8b3f42f48d153125b2de40cfb

Observation 85514468-a283-4725-a254-7017d248630d · outbound

This paper cites SpeechLM: Enhanced speech pre-training with unpaired textual data,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechLM: Enhanced speech pre-training with unpaired textual data,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.404239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.943271Z digest=sha256:d30a28f58d0254005f07ff31de273d5641fd5328a558a8d71a9056e22af0307a

Observation 86f05e72-8f1e-4a42-a720-c8b46a45e795 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.946785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.946785Z digest=sha256:678c1c1d3817644503a4959e5cbb468dba33af5e1f771f18e9ac8ef19777e8ce

Observation e74b7024-a6db-44f6-98ec-f26174eec041 · outbound

This paper cites M-adapter: Modality adaptation for end-to-end speech-to-text translation,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM M-adapter: Modality adaptation for end-to-end speech-to-text translation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.387882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.950774Z digest=sha256:30b0e16e1016a36c565bf1036fa84b697791e3f1222a9beb0a3142c218b991e2

Observation c0d740d2-39c1-4905-b8be-c793fb4494d9 · outbound

This paper cites MAESTRO: Matched speech text representa- tions through modality matching,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM MAESTRO: Matched speech text representa- tions through modality matching,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.375306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.954853Z digest=sha256:eb1f04bbe414a4e9e7ba5cb424f5998b72bcc18e2c7b3ccea53f3f3c412aa313

Observation c00317ff-d02d-48b0-9d02-c6ad41888208 · outbound

This paper cites Cjst: Ctc compressor based joint speech and text training for decoder-only asr,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Cjst: Ctc compressor based joint speech and text training for decoder-only asr,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.361980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.959632Z digest=sha256:b7ba8a1358f3365630c09fe10ae4c062b59932089468867c6b560c60646ad2c3

Observation 3b3ba2a4-aa92-49ee-829b-85887d30e4fd · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Lora: Low-rank adaptation of large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.346175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.969315Z digest=sha256:4bcf8d0159d68e6f9fb43cb66859995effb6900d95c274768ddc2adeb16e9029

Observation 7a096d0b-5c63-45e5-9e99-4ca0b139bb8c · outbound

This paper cites CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T04:44:26.028455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.964450Z digest=sha256:4aa5da6cec3043eef173a5a47f21f8f1cd1b07d1f9a0df0763796cef47172fe5

Observation 8c2d6d40-f604-424b-abfb-076846be80d0 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Conformer: Convolution-augmented transformer for speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.310290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.977083Z digest=sha256:4526e5ea71a57532e8097ff7f9b8ec179202c1b3d5cda8ad6edd268c45717748

Observation d6c1a48e-e5b1-4d07-9ecf-d18ee2837067 · outbound

This paper cites A CTC alignment-based non- autoregressive transformer for end-to-end automatic speech recognition,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM A CTC alignment-based non- autoregressive transformer for end-to-end automatic speech recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.332237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.973181Z digest=sha256:570fbcf7c473df90eabf67dc0a4c375cf3fb7973c8c8ddf4398fba52d478b6e7

Observation 490ff960-f90d-44d6-a8f0-fde17f19def0 · outbound

This paper cites Unsu- pervised cross-lingual representation learning at scale,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Unsu- pervised cross-lingual representation learning at scale,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.281353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.985859Z digest=sha256:5be3d3e6a13bd9e12431a93dfad043b24025a431ef30a590d67e786d237e57d3

Observation 4e3c3e89-b628-42c0-b2c0-3cfcd5f1b980 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Zero: Memory optimizations toward training trillion parameter models,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.980965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.980965Z digest=sha256:39ad1e0a0d5fc814c2d443687bb243decff33908c01b3711db80b8624448ae4b

Observation 7a11f7ff-e3ee-4275-81e7-b331e95a2782 · outbound

This paper cites 4552–4572.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM 4552–4572

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.537792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:44:25.886366Z digest=sha256:f50d9a1b7e204cff9bb7e1ac4f78cbfd27eb28c979a9f7a98c16a35ca69db0e5

Pith citing papers

Observation 365b54bc-1661-477d-8383-51bc92147390 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.151044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:357ca1a8ff3c6fbcb8a0350006a4c84f6e8292a258aae7cdd038c0ec49a49b17

Observation 0b04b37b-d2f0-4758-a944-eb730058e271 · inbound

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities cites this paper.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.386966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.386966Z digest=sha256:615f67aa5a5bbb5fcb5084394035154b69c5d825b93960aef9fdd8903b1ad84a