Pith. sign in

Paper Citation Record · LEDGER

Audio-Aware Large Language Models as Judges for Speaking Styles

As of 9 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 3 inbound Pith citation observations for arXiv:2506.05984.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05984 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:00.810072Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:11:44.891731Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b72ff62d-f801-4c13-82ac-941bc798fa04 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:06.258429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:58.377307Z digest=sha256:d5cef851ef7ac7d67046083560fd777ea642e0735fad4f6d2c009fa2d9d34253

Observation d61463be-9a8e-484e-b65e-0ae07283fe98 · outbound

This paper cites ### Scoring Rubric - **1**: The speech **does not follow** the required text, regardless of style.

Audio-Aware Large Language Models as Judges for Speaking Styles ### Scoring Rubric - **1**: The speech **does not follow** the required text, regardless of style

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:06.055508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:58.523659Z digest=sha256:f4a21ed640c26d9dd024c4ef3b8be3f93592053aa9f5959a1c52460f396d0567

Observation f5304c9c-175f-46b5-b7e4-580a6b182199 · outbound

This paper cites Provide a brief analysis of how the speech does or does not meet each requirement.

Audio-Aware Large Language Models as Judges for Speaking Styles Provide a brief analysis of how the speech does or does not meet each requirement

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:05.236033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:58.982187Z digest=sha256:98a1ad105f348113d984d4d17b679d921f0ed6b5dfaac4604373a5a061497392

Observation e874c124-2bb7-48d6-af13-a48197826fe3 · outbound

This paper cites Follow the scoring rubric strictly.

Audio-Aware Large Language Models as Judges for Speaking Styles Follow the scoring rubric strictly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.950118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:59.091819Z digest=sha256:f1754b5b74fb0705b6e185ecdc2159072bb1c31062873f0d33558f224ae3159c

Observation ccd62385-da1d-4c17-9c46-57a6f32ba194 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:05.798607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:58.644612Z digest=sha256:3c4b2af2ab31e7ec56f394b086379109b700f12a9354d997884465bbd070ae68

Observation f7c7cfa2-d568-4f38-b20c-a020bce6cf45 · outbound

This paper cites - If it does **not** match, assign a score of **1** and skip the style evaluation.

Audio-Aware Large Language Models as Judges for Speaking Styles - If it does **not** match, assign a score of **1** and skip the style evaluation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:05.476683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:58.788473Z digest=sha256:b152ae9d1ce8698c36dd71bba29d3f82d56c3ffd2c5b610a5ec24c0cce3ed312

Observation 355cea73-2bb8-4fe9-ae57-49e55ae65f12 · outbound

This paper cites Replace score with an integer in {1, 2, 3, 4, 5}.

Audio-Aware Large Language Models as Judges for Speaking Styles Replace score with an integer in {1, 2, 3, 4, 5}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.638300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:59.200648Z digest=sha256:838b203327fa73acdb358595e276a464d213e46d2c9b5d4331e617323b3ab112

Observation 315a90a5-f12b-4a4f-bfec-e1e15276e2b1 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:04.368926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:59.275330Z digest=sha256:3282851d268b565fc61a442749eaee21227afd303464fea25425ce978609c7b0

Observation 3a7038ea-6e87-4e53-a8dc-8f6d1ef2b457 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:04.029103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:59.503410Z digest=sha256:8a204ba0fee05de5feb9dea40e763bd9c93bf291b4098c2afa73c58bd702d45b

Observation 66f924a0-2e5f-40c7-b147-8ddb6aeeb3bc · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:03.643636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:59.734982Z digest=sha256:86e1ff41795e9ec33caa85170b720ddeb8f69c4d65cb518b392ad7d94ea1cc37

Observation 53ee33b3-88bc-4ec2-910a-99f3fcc027a8 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:03.310702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:18:59.931355Z digest=sha256:9de28f9434166231fddbdf2c504e9a6be22301f434c5932a8ccfae4e82510446

Observation 763fd7d4-6997-4b96-a147-137634dc5c93 · outbound

This paper cites Reference specific turns or moments in the dialogue to support and justify your evaluation.

Audio-Aware Large Language Models as Judges for Speaking Styles Reference specific turns or moments in the dialogue to support and justify your evaluation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.003405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:00.097191Z digest=sha256:79ad9b52fbee3606104741f9f840a8516421bdc3fb703ec3a14054236bff7d14

Observation fd0b0904-3b59-446b-a340-3f2b22d09560 · outbound

This paper cites Keep the double brackets as shown.

Audio-Aware Large Language Models as Judges for Speaking Styles Keep the double brackets as shown

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:02.577412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:00.243777Z digest=sha256:9c0511c513f7b103674480b49f0dc4a2ec06460b7342538a43793d17eef01f5a

Observation ec32e972-9c14-4504-a726-3647c47f6b1b · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:02.167402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:00.429821Z digest=sha256:906c786189860a8cd2a35e20fe85941be6a09d7e209546f9e6abd4bff99b5630

Observation 32f6549a-e9d9-43ca-ac39-87d8ce190406 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:01.861760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:00.562246Z digest=sha256:cfc7d1875d37c17ed8ea966f40e257aabca4adef54cc9b7cca8314ef9d13d51a

Observation e5a00dc6-7515-4a41-bc40-fe3e054deaa0 · outbound

This paper cites an unresolved cited work.

Audio-Aware Large Language Models as Judges for Speaking Styles Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:01.487419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:00.726044Z digest=sha256:b4ce7fc2c79397ee1a3cc81c6c497af8ead031d9a7b9fce35d5a276977f4929b

Observation 28d23887-600c-43ba-85d7-f9d69009412b · outbound

This paper cites Oh”, emphasize “exactly.

Audio-Aware Large Language Models as Judges for Speaking Styles Oh”, emphasize “exactly

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.082152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:00.810072Z digest=sha256:38b9ee4cbc7d0572d9632433d3ca69cd98f11e17d4502d20f4f6351b3c8c6ae3

Observation 532319ba-f6a3-49bf-b156-26c52dc1fe53 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio-Aware Large Language Models as Judges for Speaking Styles Qwen2-Audio Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.080329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.080329Z digest=sha256:28225b8337860fe8c763025edefc99d70a8c8261bbdc7cb87dee4bf03c57e879

Observation 4352680a-1322-4fa0-b674-cc5878118272 · outbound

This paper cites VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models.

Audio-Aware Large Language Models as Judges for Speaking Styles VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.186300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.186300Z digest=sha256:b758daca7bb8170114126d20201ec2fa6c0fec53909c659e37f873bdde016aab

Pith citing papers

Observation e329342d-cbbb-428f-8d72-7cb0cac8f484 · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.677607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:95dd45283d74cbd7de9d9a4feb7cce072224d35b7edc9000b3c6ba91eb4316b3

Observation 28fd3e27-1afb-49b1-aa3a-c8274b0c770e · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.959163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:8b9bffd94f9a5f99905769f6abf64f847b953a1b53089445ffc5da2feff6a6ba

Observation 5ef85fb7-9c1b-4e1c-a98a-de4754aa107d · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.705501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:9bb23568cdce690dd47c2cafe182f59433077298bf1b16e0f0f090c7385c0640