Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:53:26.333669Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2509.04473.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:53:26.333669Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:53:23.449877Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T13:53:27.421561Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4410ef93-6f9e-4d03-9ab9-5cba25ec4f97 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3952895-eef2-48ee-b920-228e158814c9 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 885256b4-b595-4d76-b1f4-82e21a0f4955 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings The ASR baseline benchmarks on the Librispeech dataset are derived from the study in [5], which is similar to ours but focuses solely on ASR
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 732f7e07-3f18-4a56-993b-d8257f0d5638 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings We conduct the SA training for 50 epochs with a learning rate 5 ∗ 10−4, batch size 6, and a linear decay scheduler with 3000 warm-up steps
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e497bb1-3157-479a-ae62-01fb49703981 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings The proposed model exhibits the capability to capture semantic meanings by effectively mapping speech features to text tokens that are interpretable by LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72776b61-651b-458d-8e17-569f95cb25ea · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings GPT-4 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d98b71f-fc12-4dc9-84ee-1643e5b8089a · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings A survey on speech large language models,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781a65e3-df2f-4100-925e-febc9653236d · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Robust speech recognition via large-scale weak supervision,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61053b5-c086-41dd-ac0d-63fe788e99da · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings TinyLlama: An Open-Source Small Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a2508d2-110c-4be6-afcc-a424821966ce · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23427e32-03aa-4a5c-8c6e-7bfae1829283 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c422d072-d248-469c-8148-e8c47a873d02 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bdd7516-15ca-4c7b-b5c4-84026c4896c1 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c6534f-ef50-4d68-aef2-43dc18d60de8 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de92ddd5-55ca-4b6b-a812-dc0c72dacf23 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings ” i’ve heard of you!
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 428a4dd6-40a7-4586-b460-2cf0ad19c3fb · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Whisper-slu: Ex- tending a pretrained speech-to-text transformer for low resource spoken language understanding,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50b75b74-679c-4de3-8a9b-7c016f811fb1 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings On the Evaluation of Speech Foundation Models for Spoken Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7807b942-288e-4d33-8945-a4187e73c536 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Universlu: Uni- versal spoken language understanding for diverse tasks with nat- ural language instructions,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc3f337a-3a2d-429c-af5c-37935193ffbc · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9dbc050f-12ae-42f1-a78a-099bbe3d5e2b · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Salm: Speech- augmented language model with in-context learning for speech recognition and translation,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9533b423-310a-49f5-a200-f934b58dbf2c · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings WhisperNER: Unified Open Named Entity and Speech Recognition
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3808f0ce-1201-4d6c-8054-6db1f7971421 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Chinese asr and ner improvement based on whisper fine-tuning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 388e7e80-4436-4167-a6df-e22a9740a22f · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ac977e-bbb2-4630-b908-0fa832ec4f78 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Using Large Language Model for End-to-End Chinese ASR and NER
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9bf4cc41-8be5-499a-ab66-a0aa23b369a8 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aad7e48-dc54-4584-ae1d-e9607bda13ee · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings End-to-end Named Entity Recognition from English Speech
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f73a246-f793-410c-b515-54f69c5432dd · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings End-to-end named entity and semantic concept extraction from speech,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ab6ef9c1-2eb4-42b7-a572-417e22a4c48c · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Slue: New benchmark tasks for spoken language un- derstanding evaluation on natural speech,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c5dcb25-d6a4-4dd0-a158-e93b099c9ebe · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Lib- rispeech: an asr corpus based on public domain audio books,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1d292078-0891-4d23-92c2-4d347df25a3b · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ece306-a343-488d-9b0b-d834a17db7bc · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings VoxCeleb: a large-scale speaker identification dataset
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3972282-9234-49b3-8c18-48c3e205d0be · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Ontonotes: the 90% solution,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8ada98e-7101-41aa-bc65-6b30936d052a · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings PromptNER: Prompting For Named Entity Recognition
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98be2c8-d21e-49a4-9458-0ff433743321 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Specaugment: A simple data augmentation method for automatic speech recognition,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c6d5fb-6e01-442b-b1ff-ec02188ae659 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a3bb6a-f4c9-4e8b-8947-8f08906babb4 · outbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4410ef93-6f9e-4d03-9ab9-5cba25ec4f97 · inbound
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.