Pith. sign in

Paper Citation Record · LEDGER

Audio-Aware Large Language Models as Judges for Speaking Styles

As of 18 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 5 inbound Pith citation observations for arXiv:2506.05984.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05984 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:00.810072Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:43:45.067587Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b72ff62d-f801-4c13-82ac-941bc798fa04 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:06.258429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:58.377307Z digest=sha256:5386cf830c5f678185b270e5eacd9ec0c87afab0c9d4825cce88810932929b08

Observation d61463be-9a8e-484e-b65e-0ae07283fe98 · outbound

This paper cites ### Scoring Rubric - **1**: The speech **does not follow** the required text, regardless of style.

Audio-Aware Large Language Models as Judges for Speaking Styles ### Scoring Rubric - **1**: The speech **does not follow** the required text, regardless of style

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:06.055508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:58.523659Z digest=sha256:f762925afb90fadbcc2c5601d3db8324684e0f55fff366b8d54bbe91de026d66

Observation f5304c9c-175f-46b5-b7e4-580a6b182199 · outbound

This paper cites Provide a brief analysis of how the speech does or does not meet each requirement.

Audio-Aware Large Language Models as Judges for Speaking Styles Provide a brief analysis of how the speech does or does not meet each requirement

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:05.236033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:58.982187Z digest=sha256:eb9897949dfeaef553c4750a2623af75d0efd4c3c2455af1dc2323b4da480526

Observation e874c124-2bb7-48d6-af13-a48197826fe3 · outbound

This paper cites Follow the scoring rubric strictly.

Audio-Aware Large Language Models as Judges for Speaking Styles Follow the scoring rubric strictly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.950118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:59.091819Z digest=sha256:93b663d116471ea349fb11172719274df0e062ba8f0217d56ac9fbe6b3dbc359

Observation ccd62385-da1d-4c17-9c46-57a6f32ba194 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:05.798607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:58.644612Z digest=sha256:ff1be5ff5ebdaa8ea03c614b5aa4710d2e531518535fc21a552ab9eff37f4bdf

Observation f7c7cfa2-d568-4f38-b20c-a020bce6cf45 · outbound

This paper cites - If it does **not** match, assign a score of **1** and skip the style evaluation.

Audio-Aware Large Language Models as Judges for Speaking Styles - If it does **not** match, assign a score of **1** and skip the style evaluation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:05.476683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:58.788473Z digest=sha256:cc1d5f9a3c13b328e602f84c316fe77df13a4c908502a640d0f187531b39daa6

Observation 355cea73-2bb8-4fe9-ae57-49e55ae65f12 · outbound

This paper cites Replace score with an integer in {1, 2, 3, 4, 5}.

Audio-Aware Large Language Models as Judges for Speaking Styles Replace score with an integer in {1, 2, 3, 4, 5}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.638300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:59.200648Z digest=sha256:8f1755e2639f623b4952937f765f843131e83f0bc3f61cd7646190b6075bd515

Observation 315a90a5-f12b-4a4f-bfec-e1e15276e2b1 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:04.368926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:59.275330Z digest=sha256:71ade5a044c19f7e89f54442de5da839ed3f4c8a9f7b6c808147b3e6ec5c8d70

Observation 3a7038ea-6e87-4e53-a8dc-8f6d1ef2b457 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:04.029103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:59.503410Z digest=sha256:4952348bc0bdbc85d55b6cc8f8880e5bb76f56a9499fa11b5e9bae5faf89c3bf

Observation 66f924a0-2e5f-40c7-b147-8ddb6aeeb3bc · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:03.643636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:59.734982Z digest=sha256:ad81db72c81ba5cd2c8cbcaa6e38ed2f154f18970f7b9cc22d84ca79a3317ddb

Observation 53ee33b3-88bc-4ec2-910a-99f3fcc027a8 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:03.310702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:59.931355Z digest=sha256:f3b9204bb349235b453fa01c01c5421637169ecbbdb5d8c09a75ff3d633c91a4

Observation 763fd7d4-6997-4b96-a147-137634dc5c93 · outbound

This paper cites Reference specific turns or moments in the dialogue to support and justify your evaluation.

Audio-Aware Large Language Models as Judges for Speaking Styles Reference specific turns or moments in the dialogue to support and justify your evaluation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.003405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:19:00.097191Z digest=sha256:b2242444ece5f2e30166cbcc81859065d8698786bf27b038dd61febdde55c25e

Observation fd0b0904-3b59-446b-a340-3f2b22d09560 · outbound

This paper cites Keep the double brackets as shown.

Audio-Aware Large Language Models as Judges for Speaking Styles Keep the double brackets as shown

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:02.577412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:19:00.243777Z digest=sha256:5699afde6ad4a14958498f5b0898f162d6aa1f704605e65f9bcfeef885ce3473

Observation ec32e972-9c14-4504-a726-3647c47f6b1b · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:02.167402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:19:00.429821Z digest=sha256:0b9da8c6faa93a5ab6748160634190acf1d34f40e32c14ed30fef03b6bb1c7e5

Observation 32f6549a-e9d9-43ca-ac39-87d8ce190406 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:01.861760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:19:00.562246Z digest=sha256:1474ed88cb2e489b0027adadfc3a3b3a19ba6c49f7ae1151ce153857777b2837

Observation e5a00dc6-7515-4a41-bc40-fe3e054deaa0 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:01.487419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:19:00.726044Z digest=sha256:1ac2d97394a6f5b9023eb7d5e9f8d8ebdfdc72e1b4922389ffa0c64c39171c4b

Observation 28d23887-600c-43ba-85d7-f9d69009412b · outbound

This paper cites Oh”, emphasize “exactly.

Audio-Aware Large Language Models as Judges for Speaking Styles Oh”, emphasize “exactly

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.082152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:19:00.810072Z digest=sha256:e9e9dd2981f0f5d28923765e376cf6d4346400a82d4ef429f4517c98e678d839

Observation 532319ba-f6a3-49bf-b156-26c52dc1fe53 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio-Aware Large Language Models as Judges for Speaking Styles Qwen2-Audio Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.080329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.080329Z digest=sha256:0dce4b835f720bd39b2b913b736e6a14c1f9c849a4fb89da716fbbd5d0fac0bd

Observation 4352680a-1322-4fa0-b674-cc5878118272 · outbound

This paper cites VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models.

Audio-Aware Large Language Models as Judges for Speaking Styles VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.186300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.186300Z digest=sha256:214dad282c0233d24404e9d2fe83bdce1fd71648fa03db0f7fec51e0ec0628dd

Pith citing papers

Observation e329342d-cbbb-428f-8d72-7cb0cac8f484 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.677607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:cafde3d1eb1d95b16863aa51416b97fcb9cc291606329cb43d42d5f62ec7ec84

Observation 28fd3e27-1afb-49b1-aa3a-c8274b0c770e · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.959163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:7624a8bbeac6c2432b44abeb468aac014a93d36fdca7d9fd43bd7e9af06878dc

Observation 5ef85fb7-9c1b-4e1c-a98a-de4754aa107d · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.705501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:eb310d2f778561f9da91b5353ec9309c1fdc163f8aa23e7b1d0a45baf8b6fd26

Observation d48e02a3-15e0-428b-8dde-e03c7c5153ca · inbound

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation cites this paper.

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:54.768654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:54.768654Z digest=sha256:3d34a7f0ac55c48fc61e82dc831fce76b3836d245873f247f271f82652ad3394

Observation fb5e74af-db89-40ea-9acd-822b6f268f8e · inbound

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation cites this paper.

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T04:43:45.067587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:43:45.067587Z digest=sha256:4529022cd66c0e0e8ec7f7f9c8a60e259b0e2e56551f77da6fa5ef52ee163bef