Pith. sign in

Paper Citation Record · LEDGER

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2506.01268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01268 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:05.762334Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T15:22:02.107863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.992590Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 32a3cbaa-6e3c-4ccd-8d56-d375eb2cd755 · outbound

This paper cites DeepSeek-V3 Technical Report.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.534755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.534755Z digest=sha256:f113d3461b3fd9b0660e77b58ade3a60193350236eab0b333f434ab8aaada3b1

Observation e870a575-637c-4cb0-89dd-7b34b9f23cde · outbound

This paper cites Accessed: 2025-05-.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Accessed: 2025-05-

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:06.546358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:51:04.645078Z digest=sha256:64bac57b292b02746036098385280ed8132e1364b115b7ff0a675528e5f81252

Observation f51b11e8-a651-4abc-bf39-f27f8dfcfc3c · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.785368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.785368Z digest=sha256:0e71c15eeae2a56401e6abf1ec859402b734a6f4098524bd86c1eda79608e39f

Observation 55e7fb72-ac5e-4394-8270-e0ccf9cba444 · outbound

This paper cites The Llama 3 Herd of Models.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.969756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.969756Z digest=sha256:5dada4561c69cb85def25e139b6a4ccbbbf5cea3034158a83bb382901c2095d6

Observation 6ddbaed0-c539-4b68-a36e-570db1d8fdfa · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.063654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.063654Z digest=sha256:9b74d287fbc92ff4fa67cdc07c551e7ed0fcc556e1522da8ea8079feed444ccc

Observation 1dbb0b01-1b44-4700-92fd-0b66f14ca50f · outbound

This paper cites Challenges in Building Intelligent Open-domain Dialog Systems.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Challenges in Building Intelligent Open-domain Dialog Systems

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:51:06.043515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:51:05.193145Z digest=sha256:461fe73cc01cede8ea9eb2f92a715d2ce076c2be29c8381a7a78c93680168dbc

Observation 01c5194b-dbfc-4ae5-9af7-defd59c3d951 · outbound

This paper cites GPT-4o System Card.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction GPT-4o System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.429002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.429002Z digest=sha256:db6ab60c15b8049bdb18e8fda59bd124faa016d4a9238a6716f80bc9b5f26dc8

Observation 78c27995-59fa-414b-9198-a679498649c7 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction MemGPT: Towards LLMs as Operating Systems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.571254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.571254Z digest=sha256:27694f11bccb89c4a72aab703b8cfc6643b26da3e1a511a6b8f50dd88b843cf8

Observation 8d44bdc4-d0ee-4ff9-9e65-0497dbcebbf1 · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction A-MEM: Agentic Memory for LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.762334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.762334Z digest=sha256:b3eb37431c1a270b5e960ca1cafa1a54fdc3098ed220f21da9d3b81a19ac3b4f

Observation ec703a81-7e0f-4f34-a261-8ffd45324d39 · outbound

This paper cites Language Models are Few-Shot Learners.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.320683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.320683Z digest=sha256:554d338bfa8ec975a07b1550d6ff07486ef058c1c5f697b59e27baa3126da513

Observation 513c0d1a-42c4-42dd-b0f3-cc15745d5245 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.888761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.888761Z digest=sha256:5da767816f109dedb6adfdfc3161ebd4dfd24c1fac5175384352db1639f5485d

Observation b3c97642-4e2f-49ba-b054-61c624db5296 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.676589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.676589Z digest=sha256:1ca6db6f39a48a2038a3b852677fc0c77d45d1bed31b08b8e9a0a6b9eccdba2b

Observation 901496e1-65c4-44b9-8657-027c864ba0cf · outbound

This paper cites an unresolved cited work.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:06.336818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:51:05.295653Z digest=sha256:a72a812c210b09958ae55ea8a9611a498cb395842da1bac3dbc32f6f8a243b75

Observation bdf259b0-3689-4d5b-baa4-6cf15306f100 · outbound

This paper cites Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.271321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.271321Z digest=sha256:c36810190e2839c29c9509ed3bc693f3538651205bce57dccf6bebfeb2da75d0

Observation 17274775-abc9-4f93-85c8-7927546b49ff · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.450022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.450022Z digest=sha256:f4df5450836b20504af6bfbed81ff99846771b2786538f1e5c1f32c61d076fdf

Pith citing papers

Observation da5bd0c2-fcbc-43ca-b841-79eaff1e1b73 · inbound

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction cites this paper.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.994077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:776ead80d95248ae5b97103f8ea04a26db78697e792acb05b949a7646118ce8a