Pith. sign in

Paper Citation Record · LEDGER

Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.01174.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01174 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:43.884961Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T14:53:55.794102Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2a68bdae-f164-4ca5-9802-5acdc0db4668 · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.419799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:e41728cd3ba72ec05a70cdf56fef163188cb16225b26094eb08918615b0b171a

Observation 511a8d1a-3a6c-4541-a13e-8c576cc4a2b3 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:43.884961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:43.884961Z digest=sha256:0965fd52ec2df9e584ee7a87269ea7d102e09ed7659a76bd7af69d543ccd6aac

Observation 814cd09c-f793-44f0-89df-7228590b643e · inbound

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis cites this paper.

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:25:14.922048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:25:14.922048Z digest=sha256:4299baa81f4b221c11982a75b32a77aba42ad719bdf37685eb3a0dd47a6b20fa

Observation a6902974-2fde-4daa-953f-5945ace59d2f · inbound

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling cites this paper.

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:56:57.191255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:56:57.191255Z digest=sha256:3472a2de50b406a30d0bd16d5820d217a7e05d39265e70dfeeff3a587eac00ac

Observation 84685f6f-9317-4a1b-8567-254e4f5817d4 · inbound

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models cites this paper.

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:58.971096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:58.971096Z digest=sha256:87853e9f0f99e4620bf17f116abc9970f460e58696b66a59ae16af139feeb988

Observation 70d42228-07f0-41a3-9d55-7c4603f3a435 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.647080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f62f50e5a7212be7b9b4c275c2e040de096c93d7ac5fb2afa28ad86abd53dafe

Observation e08aec35-b333-4e3c-9b1b-f2e6aa3bd14a · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.016893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:6227f5b7a4e5aa3e65a80159348dcf6a0e9a12024191492da79d4a73227a87ae

Observation 44d56285-f5f5-4bc4-b632-72ac5a7c9ad5 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.383937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:2dca86e830c8afc5fdcf8bd208780788404ad24b6810575024f835c969d64e81

Observation 04e054a2-92c9-4bde-b67b-f5b1f7662ed6 · inbound

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents cites this paper.

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.492855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T05:32:27.371597Z digest=sha256:44892243099985f7a414206af934d584727c62d36e21fbd3811d8b6028f2a445

Observation fd89cee1-7f87-4725-a7bb-0cd91dc90c52 · inbound

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents cites this paper.

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:51.432668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T05:01:44.202712Z digest=sha256:9b3bcbc1e08854413ca41abf8f8101c89b7b80665724dd44311136679c7ff6d2

Observation 0e1281f3-4fb8-42f9-8731-b258e47814ed · inbound

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine cites this paper.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.019234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:9b4e0027944b95af66243d0067a7a35424ef8608a9d0db007b494fa9b327cbbc

Observation bb848000-efb7-4446-a0e0-0df275fbda15 · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.291694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T21:22:38.888528Z digest=sha256:9723a1744bcd4b04d259bb9e8f820c5ccf7b97efb5d7145c2c8edde7bbbae954

Observation f35dd31b-bccb-46ae-a0ce-f1561604963f · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T08:53:13.408944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:53:13.408944Z digest=sha256:00a10b5009474369f237273a383d29aae2b8fdf664b781602b0e3807e38f3f8f

Observation 2a5d849a-7bfb-46ab-a75d-972749ffa3a9 · inbound

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models cites this paper.

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:53:55.795345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-07T14:53:07.512543Z digest=sha256:459db0e7164593db4dbc842c260596f06b24e496741c5e35127de6b21ab09bef