Pith. sign in

Paper Citation Record · LEDGER

URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2502.17810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17810 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:36:06.472390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db7cb67f-4973-4c2f-9e03-74dc6668fbff · inbound

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey cites this paper.

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.242696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T13:32:57.771753Z digest=sha256:a6f08d18c8dfea219cc36ad517da0970a893863d037e0479eeab3a85d9ec1fab

Observation 106fd36e-5574-4659-a943-269086e2d72c · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:59:51.100939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:edad195311f3660c86faa8c41fbd2a4f8f4ef6bc2cff410ebe75b7ac696277ec

Observation d17a3726-3a7b-4486-a695-6b0d4d77566b · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.472390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.472390Z digest=sha256:97bc2feed076b71cabf9e1068c8b6b86c9edb8545478d0b33fc369bd36dbbe79

Observation ccd32e67-ce37-47b6-92ba-2e7f39352a61 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:52:35.642915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:34d5651d53a4c619b429867fc8783353251b67110f43897b159cea1bb923fe43

Observation b60aa5e7-3069-4b7e-a59c-a4ba0bbc85f4 · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:46:03.542900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:803002422ce0e16e3e30599330cb543b734d0b5a9536470f69874c1fb277d205

Observation aba706f3-94e5-400c-b97d-419f63c5aace · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.937202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.937202Z digest=sha256:54b989feae9048cadeaccf9b1ef4ccf381a8762cc31af81761fd81d80f7929eb

Observation c3c10ec2-4a90-4f64-8d47-f4c4f0029954 · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.504670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:49e5d93cb597e669b104655fc301f97243a84a83a63ae9cc5eaaf05d97bdb2cc

Observation 5be8038e-bcc9-4550-ba44-09d2b089ce63 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.186750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:ec3e4a1b257b151707f96d3e0a23393992433cda89fcb2e7fc63098b61e69df2

Observation cab538ae-c1ca-426a-bfce-4a2d31dfbdcf · inbound

Evaluating the Expressive Appropriateness of Speech in Rich Contexts cites this paper.

Evaluating the Expressive Appropriateness of Speech in Rich Contexts URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:56:14.573707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:56:11.473572Z digest=sha256:6a05b9317a1b68e2a226a2db802337ad63a7ff7300613203032f3aa9c1388f7a

Observation 47f6c4d8-2d33-41ab-8c32-a3db3e5085b2 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:49.047504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2b85f1676cd9b888c5d782d63bd2e1187f2223f43e2329d5098dbcc5e2067073

Observation 330ba7bf-eb71-4b78-873f-bddb8994d002 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:55.060576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:eca7352f1f5c69b4b93524b3122fb965c4432ea965bf90849d87d26e63ae6915

Observation d981c1f0-b2bf-43c5-9c63-56797b1abc01 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.301296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:dd540da193f88739ae98f7723dab648621b42c64f12ce7ebe42813d797fc747b

Observation bdd1c086-e010-41c2-8fd5-4207109c55dd · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 129

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:09:24.437109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:f2127b39bca2b82be1a1cd1ad5532448ae6d9f4f155045bf48e2a27bea02daad

Observation 4cad0548-0107-41e4-a95f-f154e1f5ac0e · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:27.009937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:08ad00fa5108da8b6fd58fb7bb1a1df2bfc85ff624c95a518e51a7748bd80d27

Observation b94e9da7-36be-463a-aa74-4ad1a086ea3d · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.235234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:e5a10ede12cae824cd4935b82dcdd9d07223268751f9a88e179732ec7b6b3413

Observation 30a4e08a-3965-4cca-931a-e968a2deb774 · inbound

Liberating LLM Capabilities in Full-Duplex Speech Models cites this paper.

Liberating LLM Capabilities in Full-Duplex Speech Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:15:09.030108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:13:14.701980Z digest=sha256:f3f545568eaff700cdb5f083e374e9f8673dc83546a04b8d34730186500dfa0f

Observation 533cc13a-9506-4390-bb4e-c9a1d0468fef · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:18.079021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:18.079021Z digest=sha256:d7f44849a721473ba353b34bf3fb13a67898636c3e53e8286b34f3c3d77ca62c