Pith. sign in

Paper Citation Record · LEDGER

WavLLM: Towards Robust and Adaptive Speech Large Language Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2404.00656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.00656 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:49.748472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:37:34.396958Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 800bb696-9382-4988-94b7-9b2b9178d651 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.375568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.375568Z digest=sha256:80feee060e4b06ffd017cc77610ab7846acdeda372154681f5ee0164f4645905

Observation 0111d5a5-52f0-4b14-b57a-99782459c5a0 · inbound

BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection cites this paper.

BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:37:59.256646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:37:59.256646Z digest=sha256:ebd770f4d0d1ec4714aefed0fadec31f1c99e38ac62510c6a206f5c64c59058d

Observation 8a117a4c-2043-4e70-b042-265b3e9255b6 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.849435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.849435Z digest=sha256:a8926a0148adce35c6bce7ba5fbb51c03ab0c9ba4b165b1d3318a8486b0271d9

Observation d1ce8f71-f084-45bb-be76-1c4d9fe0a218 · inbound

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models cites this paper.

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:40:25.437559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:40:25.437559Z digest=sha256:9349c02b2a15ef3b340f10f33cebf6cda63927a89e31585324633bbcfc44cb9a

Observation df371e8f-507e-43ae-8b4e-3424983f3466 · inbound

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison cites this paper.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.599424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.599424Z digest=sha256:9504df6af5af007ac9e2a5b8fb3797c056c86ccad15fc49a08f7eedf0a2c9cfc

Observation 5bad9113-8d95-4f27-a1a3-14cd052698d5 · inbound

LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations cites this paper.

LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:17:35.497624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:17:35.497624Z digest=sha256:6ca6b87f3e920fdae15a12f3c944ec2c7939d07ae7200256d4141c98243cd09c

Observation 46b82778-5d37-4368-9455-c69442a9b454 · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.781684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.781684Z digest=sha256:1f757f72d4124cde46402fe350a8a4fc0f53979afe5abc8da24f673f210a8a25

Observation 1ec89678-f4b5-4d0e-a619-1e8f71e72fff · inbound

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators cites this paper.

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T12:30:52.072980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:30:52.072980Z digest=sha256:27d65b46f90df127d8dd67f87c100f9a837dcd8fcb1b1e944df835456d8d2410

Observation b680d517-637c-45a4-abb2-09fcb9126e8c · inbound

A Preliminary Exploration with GPT-4o Voice Mode cites this paper.

A Preliminary Exploration with GPT-4o Voice Mode WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.327669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.327669Z digest=sha256:8b14465ef5c2fc68f0d5c1cb0a988c703add36e6dd100c717a379c11446ced91

Observation eb985253-43d0-4221-9e6f-15dd2dab079f · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.359286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:8781607e05539517aa2e3cc300321f0cd128d311df6e7fea3dfceaaff8d6d584

Observation a302002f-5c66-46e8-87fa-14ce7f01c409 · inbound

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits cites this paper.

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:49.525285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:49.525285Z digest=sha256:6c50c17af918cc74facad8de86c729c5e11085151a7b90bccd1af47739f23707

Observation 23f2d7f6-e5e9-40ff-8a66-6420ed41fd2b · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:47.276785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:47.276785Z digest=sha256:ddfac9f3e3dc7db49470eaf46b16a8334497d182d253c3528d2847ca279ef875

Observation 39dc1cdd-a81c-44c2-89b0-8999099d6d37 · inbound

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models cites this paper.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.067293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.067293Z digest=sha256:8e14b08123636b247269870ef688d61ae8a3e864176b169606f7acdf0280ffb0

Observation 9758e2b8-382e-4345-be9b-0970a5209c8d · inbound

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs cites this paper.

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:33.844445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:33.844445Z digest=sha256:0da7e9036aeb496fc87fa600cc997e1df2be332fa3b446cea9473fd5d2ea46e9

Observation 0f9d2bf0-68e8-41e4-8fa5-9059d48c1758 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.743399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.743399Z digest=sha256:e34a2f9945b87069cef1e6ab14a8931a2b81adc60ba1407002decd340cf7cb0d

Observation c481c263-ae91-4edd-b67e-aece03537254 · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.593442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.593442Z digest=sha256:5c01d889d46122f472e7806b1cf66833574644b14557a5937416017ce5bfd01a

Observation 2d50605a-97c0-47c4-8eaf-39bcf667dc6b · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:30.469502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:30.469502Z digest=sha256:b108d8562bad22602b1800814ba61e93d1db8c4b3756ef97283359c66b7ff18f

Observation 0901f8d5-522a-4fc3-9a6d-41b30ebce6f4 · inbound

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition cites this paper.

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:59.761074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:59.761074Z digest=sha256:0a7ddf5a6d071ab4f2c38b04e7e52f2989cb0c6154ac772bb641541ae915abd3

Observation e44347e4-7050-40ad-adc0-08546cffa511 · inbound

NAVER LABS Europe Submission to the Instruction-following Track cites this paper.

NAVER LABS Europe Submission to the Instruction-following Track WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:06.539254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:06.539254Z digest=sha256:aa779d7bbeedb1c11abe33094ed8f6202bcfd464ec1b535f697ecde1ed83bd88

Observation 30911dfa-cd1f-4296-b876-2eec645a6286 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.527166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.527166Z digest=sha256:6a23a135185de6d2a7eef71159decf33623cf1f034cc3e37bb9d8aeaf43158d5

Observation 4fccb077-0b67-487e-996f-acac6bec9f86 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.712860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.712860Z digest=sha256:f37b773fa55be03145a945d77e8a0c9c142603e519d152f5a94a28754474bca3

Observation 0fdf5d1f-b198-4d7e-9318-0b7fc8e144fa · inbound

Self-Improvement for Audio Large Language Model using Unlabeled Speech cites this paper.

Self-Improvement for Audio Large Language Model using Unlabeled Speech WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:49.748472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:49.748472Z digest=sha256:fff03399df59446c25da2b16299d2139c5b9a0cfb90a4574e0c41ba5456f5259

Observation eddac622-e3c2-4186-a8bb-b36bb8e41ac1 · inbound

A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR cites this paper.

A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:03:14.524557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:03:14.524557Z digest=sha256:51a6d6d3dfb860876f29f2932c9721e9f0fb3728ae741c176968f26b6201b5c5

Observation 3a65ff3c-010c-48c3-a432-af3a74704ccc · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.512556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:b9a7763947ac0ebde26cb7a57dab8adecd54740a9ac9b3d3ac4039b051811a07

Observation 9bf4cc41-8be5-499a-ab66-a0aa23b369a8 · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:24.835833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:24.835833Z digest=sha256:5725fa365ddc6bbcb9cd25081369b6cfb832de1dfe2fef38103e3a94e2c2abe4

Observation 972007dd-402b-4d7e-aa27-778be2800211 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.804002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:d6bd1940a2efa077e94226cc03048d8cda6cb347ab3dd33888772104a344e2b1

Observation 2dab21e0-b00b-4940-96c8-083110bdb8a5 · inbound

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval cites this paper.

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:39:58.380231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T02:49:35.819155Z digest=sha256:b20829f2a8978522c906acaf140bb0f4d4da120c947701a8e914550a46132bcf

Observation 8e2f9176-2dbf-40c1-99b9-f4b58f4d0be4 · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.183604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:769476903e68dee19d868bfe7c3ae401a144f43444db1d2721b8b2dff583cae7

Observation c774a0a5-b0d9-4324-803a-1a86053f725e · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:50.946446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:50.946446Z digest=sha256:5225d04a94415ceca9911ccb60c392cc05109e167996d147dad056777361d228

Observation 63af21e0-5f45-421c-b25e-25dcd9573349 · inbound

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track cites this paper.

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:08:32.423186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T15:04:02.640015Z digest=sha256:9e0f28219d3a9fe7a6538730ace07fc1d2f4cf3fc777a4a1757fd7c4589fb10f

Observation 6e4c7289-d80d-4a0b-906f-af201dded00d · inbound

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models cites this paper.

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:53:55.789976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-07T14:53:07.512543Z digest=sha256:e8efdd1098e75e0e8baad2f1ee8b005922e3371a00ed338024ec16df0a98582f

Observation f1d48f1e-066c-4610-8783-4b66cd7f5a09 · inbound

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task cites this paper.

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T04:47:56.963554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T04:47:56.963554Z digest=sha256:f7500d30dc6ec6f4660f959d8013957249206c28f9e75541d71fc5b8ae4d820d

Observation afeac008-e22d-4945-a197-29c902ed6a3c · inbound

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task cites this paper.

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:30:02.391299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:30:02.391299Z digest=sha256:8149bfad2368ed5216915bd44acb3eee68fd0269dc6f9494bc1e1c3ca8c1b66c

Observation 3fe8e6f2-f109-49ef-bf16-6fcddfd29a9d · inbound

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs cites this paper.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T20:37:34.398583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:66bff8e17550f896e3e713390a686a60122f2eb0ab7f172461d23de86c9a6916

Observation 9d169afa-8bba-4e92-95e4-1b1b83d6a072 · inbound

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models cites this paper.

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:18:27.114855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:18:27.114855Z digest=sha256:8577b889f1475e0cda831ac6c6ec8133a34a5eef8b820d187cff5570eb07d04c