Pith. sign in

Paper Citation Record · LEDGER

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2506.02457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02457 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:23.320594Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:22.389818Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T07:39:48.680240Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24a79a7e-e7f3-46b6-8e29-f2eabedc2326 · outbound

This paper cites SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.389818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.389818Z digest=sha256:0034115ba8fce44efdd3dfd699b00f38314b8ca05ac4951a3b231843d5c6767e

Observation 6935127e-fa43-4f0d-b77e-26f4d1a3a331 · outbound

This paper cites Speech LLM Speech LLM extends the understanding capability to speech flow, performing modality alignment between speech and text via an encoder with adaptors.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Speech LLM Speech LLM extends the understanding capability to speech flow, performing modality alignment between speech and text via an encoder with adaptors

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:26.828263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:22.451786Z digest=sha256:35a9679cc5429857656aca6c60180818bae442db01c74eafcfb2dea1f310ec19

Observation 335ffbc3-ee3e-4578-a789-24643a2af936 · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:26.608938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:22.535835Z digest=sha256:b4c631858a1d3e511a0e4f88ad7fa099f47ccc2d1fcc8c79586b701c5a62c159

Observation 3007e468-ce9c-436f-8f31-d8c14bf5e103 · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:26:23.863004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:22.625987Z digest=sha256:771cb49b2519095c972b9e2544f4a5755f6a3f068737b25fc89f3fa1393494af

Observation 741100b8-5df6-4d9e-8de8-ec39610b01af · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:26.436653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:22.693809Z digest=sha256:dff069d1ff604d6ae950a2754b2de0277b661cedd9428a759d97c730950c8577

Observation eed943fb-c74a-42e8-86bb-cac7be6255bb · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:26.295669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:22.770624Z digest=sha256:7f85a6e18edfed748b11bbcebf2ddbcb6cdd8caa59c0f9f718352cf5ae6c9fec

Observation 04f99b43-0c47-40cd-b373-185506acfab5 · outbound

This paper cites GPT-4o System Card.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.870915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.870915Z digest=sha256:540b7b6c69929f1c4e8d71e84f87988e4b98f34e31448faee95290a83931dd75

Observation b1ef39e5-7134-495b-aee8-c300409e6154 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.926836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.926836Z digest=sha256:e8c31ba525203812de58959e7176f490c85c3219072ad5361ddb0a783b402dc8

Observation 3a38c437-6768-4ac8-b6b1-defd92aa8332 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.008363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.008363Z digest=sha256:80e64f7081de2c35097a6da5a909259585692b16c45358367810e6e79e39ee51

Observation f821b026-4326-4f47-b148-bb854874ed7f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Moshi: a speech-text foundation model for real-time dialogue

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.050878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.050878Z digest=sha256:1e3c85eeeb5dc186a4357f3b0ace61116e6a9296bbc0e7652782c67364824504

Observation 657ac16c-0126-4684-af44-6e56729f7170 · outbound

This paper cites Dynamic- SUPERB: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Dynamic- SUPERB: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:26.163525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.138843Z digest=sha256:79c5a9cc47c591fa4338c761233e7c06f89675f9d2e38203d096bbd962834dbd

Observation 5aad3ced-54b1-4a76-8309-61beb49207e2 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.197926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.197926Z digest=sha256:ccb73e60c57525a1bb3be8335432f024640a38a87d89e762921694d853d9021d

Observation 0db694d5-841a-41cf-99e1-1c9efe1a8c85 · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.203188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.203188Z digest=sha256:471d410be1855d1b4cd1d98eb2a14a7e99e8536f7bccfb22c355db6adaa12df1

Observation 1dae2d60-7300-4edf-ad27-3acb7069dff2 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.208362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.208362Z digest=sha256:8182a54422080632db756ed2fab3a434d8bb45b85ff1c19a962b535ec861af76

Observation 277382e1-99a2-4fc9-ab1b-e2432272bb66 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.213004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.213004Z digest=sha256:3b46aa115236def4cda7a5184a20748f16f0474b591575fce9b593c784ecfffd

Observation 3208b10f-2a13-4b05-85bf-ef46656c5d91 · outbound

This paper cites SALMONN:Towards generic hearing abilities for large language models,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SALMONN:Towards generic hearing abilities for large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:26.018205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.217381Z digest=sha256:346dbf567a351d83421c99e827fa9b62cd867be55c708039320e12d1cd2b55e1

Observation c243da26-5a10-41da-9c47-b0fc12b019fb · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.916923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.222170Z digest=sha256:a68c0fc353ccd4e85963be81a5d73a17c0359a7ae7a2baed63a8250219492b59

Observation 2f99f1ab-e977-4b39-b9f2-dc13e6840bf7 · outbound

This paper cites Qwen2-Audio Technical Report.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Qwen2-Audio Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.226286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.226286Z digest=sha256:de96a9d42f97a626e5d4866a06d0fddbf30e3caf18094e7a2c3b86c77ad56e45

Observation d72b60f0-b1c0-4ad4-86ed-f944745a5b04 · outbound

This paper cites SNAC: Multi- scale neural audio codec,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SNAC: Multi- scale neural audio codec,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.756587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.230921Z digest=sha256:2d13e9d9989a848ff3f82a4975233c59961cd1769e241dd0da3f4b689427d72b

Observation 78bf5a99-e0a4-4aac-ba3e-64c54fd2a8ab · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.234959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.234959Z digest=sha256:4ffa3f6bc9bd4342722d087bb479965e63d46af2687b6a51b8c3c57a83b44464

Observation 25d6d93a-a35b-43e3-a76c-698b211e4008 · outbound

This paper cites HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.610697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.239575Z digest=sha256:a3116f3f8eedea7e96f67531737642c02058fe074f466e4ea11acaf39d4a5c65

Observation 81babd28-831b-43ef-be8d-cef09381c961 · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Speech resynthesis from discrete disentangled self-supervised representations,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.430775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.244104Z digest=sha256:9179819d0ce47dec2918160fd0395552d1929f8083c0f1cb272004698082f3a4

Observation 37d13974-9f50-4609-bf99-3a907d1da2c4 · outbound

This paper cites Westlake-Omni,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Westlake-Omni,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.265736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.248523Z digest=sha256:f5133f1ac3f007aa7e26b70cd2a19a6afa4fd86ec84de69fa37800f7f43409f9

Observation fd4d79c4-b7f4-466a-bb6d-a80d9ec59e96 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.253202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.253202Z digest=sha256:f4dbbd5ab4491e7e0cdb5607711d4a4be38b503ef86b7e2af7b1021cfd9583a3

Observation f725328a-7a76-49dc-a579-cb24b7ecc36c · outbound

This paper cites Be- yond turn-based interfaces: Synchronous llms as full-duplex dia- logue agents,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Be- yond turn-based interfaces: Synchronous llms as full-duplex dia- logue agents,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.096745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.257541Z digest=sha256:6f3814d22a670b39ce74e8f52a67c7374464ea50685d0cbbc6b857ed2e6e6e00

Observation 101c2dd5-5392-419c-820a-30b3c3d19b08 · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.261871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.261871Z digest=sha256:ce9ec5e964106cc3c48514770683ee185199c8cf6c8da7e47ef6545513f46fff

Observation b577a389-a753-4102-adbf-4c98624ca753 · outbound

This paper cites Baichuan-Omni-1.5 technical report,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Baichuan-Omni-1.5 technical report,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.266830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.266830Z digest=sha256:3bc7dc0f590d5524b906684dbdaf9cbc4a5216922042c22adc39aebe48463297

Observation f9735e47-eb7d-479c-a8ac-03dc3d0f8cf8 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.898400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.271445Z digest=sha256:dad5752b06423553f01903e0ad67339c5dee8d2583ee15fa0192c425da3b9ce5

Observation 2733f0ee-bbe7-44dd-b717-497f1ce9d094 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Lib- rispeech: an asr corpus based on public domain audio books,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.276005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.276005Z digest=sha256:70d2beeb63760c909aa0b21ab7b563b034610ae52e98b2906486ad4c436506b8

Observation ffa586ec-7787-4748-bc32-a3b4aa3794d3 · outbound

This paper cites LibriSQA: A novel dataset and framework for spoken question answering with large language models,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LibriSQA: A novel dataset and framework for spoken question answering with large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.706381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.280776Z digest=sha256:57be3d639f0f1551e63ab846ba9757e5a129c2495e8624f6f55ae2a2948ba549

Observation 4ad395e0-aac3-43e9-820e-feafba8e888d · outbound

This paper cites Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.560476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.285411Z digest=sha256:d6feb4b1f7e9a55c006dffb3f828d6cd995108097c67626d5e50aec6bcb854c2

Observation 6d061c8f-99cc-4c34-97b1-977e6766fafb · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant IEMOCAP: Interactive emotional dyadic motion capture database,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.289433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.289433Z digest=sha256:266f4e5b524614aabd6bfa610683e493e82e4f0a4d6dbb2d43cbdbde7211facd

Observation 3f915456-202c-420d-b2bd-e2b094f68824 · outbound

This paper cites Common V oice: A massively-multilingual speech corpus,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Common V oice: A massively-multilingual speech corpus,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.374900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.294124Z digest=sha256:15e47a390b97f0dc81320c32df0d4204fccf4799e7387981949eb1466e672ea1

Observation 16d4b01d-6eed-45ce-aa65-308b97d668b2 · outbound

This paper cites Stanford alpaca: an instruction- following llama model (2023),.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Stanford alpaca: an instruction- following llama model (2023),

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.140936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:26:23.298170Z digest=sha256:2af86ee4a7e571dd6236ade92e9bccfa17da772c3f72dbe33b4cb0c7a9e9c7ef

Observation 3ca2bfd7-9187-4d2a-945d-c5e9210fa9d0 · outbound

This paper cites The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.302596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.302596Z digest=sha256:c2203ae5bb2fff9f91f62b1714b1eee47f9de7ecbc0cb22299140160f6040bfb

Observation f82a96c8-3bec-43ff-899d-1ad5b1167973 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.307131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.307131Z digest=sha256:602924099a8ca20fd10885f6b0ab5af657a134f865f0900e217633b2a1832870

Observation 3141e860-23e4-4cdc-b2df-98ae5a89919b · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Robust speech recognition via large-scale weak supervision,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.311743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.311743Z digest=sha256:605f3ef5b179825b2ebfc9003ad998de7cdaddd451cf917b42e23d9c462987b0

Observation 4b7c8061-2bde-426b-8c21-ff56e53f4e71 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.315974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.315974Z digest=sha256:d4fd164bc53a87ffe482deae4098b59d28bb9e0f5554c67ea226f104dafcdfaf

Observation 8b15efbd-c7c1-496d-b92e-9470a0a2afa2 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.320594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.320594Z digest=sha256:175ebc0c709997b1f12748bdbe7631005f93ced7d181a3f4a35f955d7e0ac787

Pith citing papers

Observation 24a79a7e-e7f3-46b6-8e29-f2eabedc2326 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.389818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.389818Z digest=sha256:0034115ba8fce44efdd3dfd699b00f38314b8ca05ac4951a3b231843d5c6767e

Observation 77213481-1a76-4b50-ab3d-5c2d1777249c · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.681853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:8c9f173780f70c193a2effe3e1cbf7e5162c13b47551b7fd04171e85826bf9a3