Pith. sign in

Paper Citation Record · LEDGER

SpeechVerse: A Large-scale Generalizable Audio Language Model

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2405.08295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.08295 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:49.813356Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 474d7e28-d8cc-4621-ae66-e608e21415dd · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:14:45.462747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:a093701fc91caa6669ff4c282360b4d79f1a4792687ea53587c6e9daf9d74cd9

Observation 4d1dd1b5-327b-4a08-8f2b-1ff04dd74303 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.160614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.160614Z digest=sha256:6fae6aad285de294bb6564b561b8ef8e335e49e62daa086ad08b0907cb07a110

Observation 7df69f19-523c-44be-8936-bd2a9a07ecf9 · inbound

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM cites this paper.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.866877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.866877Z digest=sha256:71a96a9cfa6794a0fb9fe2f0501ce5269606a5a161aa2757320b619cd8139fc6

Observation 07045fb6-57ee-4793-9e2e-aee6c25e3365 · inbound

SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval cites this paper.

SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:27:27.306355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:27:27.306355Z digest=sha256:78df83f3f0bdc1d62de7bb0b343b6944ef8f1a9edfe776ec153b4d3c654714b7

Observation 4354f0ba-dcae-4d10-b5ba-6989fe90e10c · inbound

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks cites this paper.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.616757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.616757Z digest=sha256:3df1845b9f9890fbefd6f5e417872f14d144f7fa86c7d4663063b8f0a7ad0235

Observation 0137b1e5-542f-4dbf-a1c7-5b910742c6b6 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.455567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.455567Z digest=sha256:bdec8a99ec5c81b67a639b52792310a4b01841fababd1ce07e7ea0137e278284

Observation 986c61e3-f5ce-4887-9e99-81673de043b5 · inbound

A Non-autoregressive Model for Joint STT and TTS cites this paper.

A Non-autoregressive Model for Joint STT and TTS SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.366374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.366374Z digest=sha256:9fadcc29f8ce683581974297fa71079f0602c9c04c4f73e9c7dec9668a2ee6dd

Observation 03848df2-7aae-4738-9530-917405104a4c · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.683344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.683344Z digest=sha256:31aea497418df5424486983dbc0f76b1b3cc104804d565a26a8aaa4415d66350

Observation 698c57b9-b00b-4d48-bd6b-cfd0734fdf31 · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.652272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.652272Z digest=sha256:02ccca03769d51cbfcb9dea6c5dd2f5faf77e80209f41c5792d9f95adf514ec0

Observation 906cd69b-d30e-43f9-9143-7bb3626278ef · inbound

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators cites this paper.

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-10T12:30:52.057019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:30:52.057019Z digest=sha256:839cd47b86add1e0b83fe4edfb4605b37d4f4508099d202136de4fa6b218ae2a

Observation 36dc9981-931a-4677-b49c-686fae2f3684 · inbound

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey cites this paper.

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-10T04:36:38.420849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:36:38.420849Z digest=sha256:232073a29ec07d03cb6f4750b7fd62f42c470119c419e4026a33280dc17334e7

Observation 68bbb0fc-c561-44ac-a269-d8ae6f1a0680 · inbound

SparQLe: Speech Queries to Text Translation Through LLMs cites this paper.

SparQLe: Speech Queries to Text Translation Through LLMs SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T22:05:23.794304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:05:23.794304Z digest=sha256:71087a7376190b454eaf5ca3e381df992431ae7a89d291ae3604935e88e8d0b4

Observation 62792d3a-d3f9-48b5-badd-2459f82d3802 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.291681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:65d059ed86dfaf6718b49f7fac7e26bfbfd5f97d281225a1fcec51aa04970127

Observation 6f095213-9732-49f6-9f03-231262498176 · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:54:03.293192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:90ff9d438929bf6c8a6f657e378a0d6a7dea39669b14dc8324e9d0e0549d3f80

Observation c9bfbf4a-251a-4423-a3d4-c7be0a6c46ad · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:07.987371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:22457a56ebf169d64e1489af8171db910b5174dabd7b1e96485969ce4b0b7740

Observation 1e9316b4-b1fd-426f-8937-3caf555a31f6 · inbound

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving cites this paper.

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:32.222856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:32.222856Z digest=sha256:0f2711622a079c7f46ff4139be234eef03a24d63f7c9a8f7851b1d614caa7515

Observation 443ccc8e-7d58-43a8-ba77-17d99a880be4 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.293820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:8eddf59d56b665bc1c341ec51c776f0b0bf6bc901525cc1278feeadacc5d0cb1

Observation 0a06db11-99ae-49f8-8285-e4a32689457d · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.082340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.082340Z digest=sha256:83c87fa2ba6f0c8657c89740474eed4aeed62a99f0c07c7703debd469e456421

Observation fdb3b362-ebbd-455e-8dd6-55c2a7f4ad4f · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:31.841304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:31.841304Z digest=sha256:aebe0d17184765ce2427ae8be6588777bb816a104121e8a2d9918823528f6d2d

Observation 8c41e6ba-b84d-4f4b-93a4-d03b4893ffbb · inbound

Self-Improvement for Audio Large Language Model using Unlabeled Speech cites this paper.

Self-Improvement for Audio Large Language Model using Unlabeled Speech SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:49.813356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:49.813356Z digest=sha256:4c45e1dd108a4f6acadc212b75ead66d09098dd75a356a1d2743c73d46fc498c

Observation b91d2310-470c-4a35-a5e0-5ddf7a956343 · inbound

Your Spending Needs Attention: Modeling Financial Habits with Transformers cites this paper.

Your Spending Needs Attention: Modeling Financial Habits with Transformers SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:59:03.105773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:59:03.105773Z digest=sha256:be665680f667027e17cba7ee271752f68b84004c7e16156b07526d3dcca029be

Observation 89172613-8aed-4848-822a-cc168751ec30 · inbound

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition cites this paper.

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:01:21.059534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:01:21.059534Z digest=sha256:9259b3e2c2275575b55646d48cb9c13246fd9583e691273c2be76e55a4514c8d

Observation f54078c4-3e0f-4f78-99d6-2b2fa9376b3c · inbound

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation cites this paper.

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:27:29.283691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:27:29.283691Z digest=sha256:2a127be289d9a814f643f7edec7292fa52d94ed05784dffd79da8319aeffb7ca

Observation 44ed0d62-41c9-427e-9093-19ea7800fef0 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.516544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:a0d30642ea0d17ed9258a48abd54bef102c45207b82823f26d71bc99ac7b44f6

Observation 24cce4b3-79af-445d-9dcd-fb45577f3dec · inbound

Direct Simultaneous Translation Activation for Large Audio-Language Models cites this paper.

Direct Simultaneous Translation Activation for Large Audio-Language Models SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T16:31:36.857609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:754bb16248c99a7c200adec1f69569e11de3e47660be0349eebb8cdd511aeb03

Observation 9a25b118-325e-4ad3-a8db-c625e6a19e72 · inbound

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification cites this paper.

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:48:45.994945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T00:46:08.921194Z digest=sha256:b4693d0779d93dfd24e4e92d1b533eba3eea11696ad15836f70d9b2b4ff15ad4

Observation 7950c080-4db3-4b22-9a33-64b1ead58181 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:37.503517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:37.503517Z digest=sha256:2539618f1333910aa897e82464ac796139eb08237f245a59d79c14e3571ce974

Observation f701f70d-084b-4191-93fb-9320c4c6b1fa · inbound

Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection cites this paper.

Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:40:51.416440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T10:37:56.416600Z digest=sha256:2ed3873953d1cfc5e9620b0460ba54686dded4a5a436bd972b9e28da9c50aabd

Observation 9e4419e8-9bf8-4396-862e-f2a4021a6b9c · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:50:40.260650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:29c759f35c9c0bcb82ebd1bb2380a082e51a1f263fab12ceb094638cb6f6a80b

Observation bb44bd26-3154-4e60-aea1-c6402245406d · inbound

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity cites this paper.

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:50.813861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:50.813861Z digest=sha256:e8a833758ac755a403da0e0f1bcd99dcae5e94c8b519ef3dad0644a2e17849f5

Observation 8875ffd5-8b02-434c-9207-648105f67899 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.282925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:b7f74ac6b4cfc037ed5dae6ade5d439fa1b4c1e0395c5ab664123ca420ad967a

Observation bcc2be79-4c71-42fa-8a17-0086fe8db9dc · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:48.981456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:78d39a292ccb2bf51cdc92a83ba7cbe00151e7d6ce55b3f1cbf18bdc15f584b3

Observation 60960240-19b8-47ba-b252-6a48a7d72167 · inbound

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding cites this paper.

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:54:45.172403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T06:54:41.345212Z digest=sha256:e73926639e7651f0503887b4ecd8fea18ca7bed8e908e925471b430a867629c3

Observation ed66c343-5d3e-4173-8a20-25e1c6f7bd57 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.347302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:4da009109b2adc8250fb8a6ad32cfe4005c9dad13ad62d68be5241e285e9e035

Observation 6754d537-dcf4-4b2d-b544-b7a8c614195d · inbound

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition cites this paper.

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:29.122979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T06:51:41.995108Z digest=sha256:c2e28ebfd7aec00c7411b5d6238bcb11456a84cf9e0dd133eff4dbf938b782fa