Pith. sign in

Paper Citation Record · LEDGER

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2506.01268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01268 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:05.762334Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T15:22:02.107863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.992590Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 32a3cbaa-6e3c-4ccd-8d56-d375eb2cd755 · outbound

This paper cites DeepSeek-V3 Technical Report.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.534755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.534755Z digest=sha256:8db8ec543cc6fbd117d230d574f317741ca23830024652cc083247cf58a88d1d

Observation e870a575-637c-4cb0-89dd-7b34b9f23cde · outbound

This paper cites Accessed: 2025-05-.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Accessed: 2025-05-

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:06.546358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:04.645078Z digest=sha256:5b1261e72cdd325aefcc206a50024f4cfb8b7a2a95133124cb787ad431ab28f0

Observation f51b11e8-a651-4abc-bf39-f27f8dfcfc3c · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.785368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.785368Z digest=sha256:b9b42499e65ae8f33627da43e9017cf825d5335c22773271b61f507497a430ba

Observation 55e7fb72-ac5e-4394-8270-e0ccf9cba444 · outbound

This paper cites The Llama 3 Herd of Models.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.969756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.969756Z digest=sha256:329e8f6dfa8891cc6459503833a1099587558a11f75b6e63fd802d1d475d0697

Observation 6ddbaed0-c539-4b68-a36e-570db1d8fdfa · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.063654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.063654Z digest=sha256:87eafe4440b851d0d4a7e54ad974bdec59fe20d5eada23d25da43bcdb0dd8c0d

Observation 1dbb0b01-1b44-4700-92fd-0b66f14ca50f · outbound

This paper cites Challenges in Building Intelligent Open-domain Dialog Systems.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Challenges in Building Intelligent Open-domain Dialog Systems

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:51:06.043515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:05.193145Z digest=sha256:858958e686ec3cb02e10f2ab9315195d4e455d744518e13a530e0811ac733d1c

Observation 01c5194b-dbfc-4ae5-9af7-defd59c3d951 · outbound

This paper cites GPT-4o System Card.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction GPT-4o System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.429002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.429002Z digest=sha256:27e88f91e7a4a1c3f343d50f1d1f8fab7f6f19dc51c11e6c98f5001a345aaeac

Observation 78c27995-59fa-414b-9198-a679498649c7 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction MemGPT: Towards LLMs as Operating Systems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.571254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.571254Z digest=sha256:19853ac215fd3eb0643959de7659b82013ae9d2c7e8e14bb1f249a3d9c2ffdfb

Observation 8d44bdc4-d0ee-4ff9-9e65-0497dbcebbf1 · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction A-MEM: Agentic Memory for LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.762334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.762334Z digest=sha256:f63a0169106fc7716b22dd14a5c7a60e3bd2091515c3a8805a414368334e91af

Observation ec703a81-7e0f-4f34-a261-8ffd45324d39 · outbound

This paper cites Language Models are Few-Shot Learners.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.320683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.320683Z digest=sha256:2735223909ed69acc9f93c9dc39b4c4bca371288f5be46915c80131f3116c2bc

Observation 513c0d1a-42c4-42dd-b0f3-cc15745d5245 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.888761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.888761Z digest=sha256:97863ad69ad5d6c20a58e0ba002d736d49a99a49b1899cb45381888a2d4badce

Observation b3c97642-4e2f-49ba-b054-61c624db5296 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.676589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.676589Z digest=sha256:3caf99c3ee2005433438ff130a745771ca006d39fd65140f7d244b740a3594a4

Observation 901496e1-65c4-44b9-8657-027c864ba0cf · outbound

This paper cites an unresolved cited work.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:06.336818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:51:05.295653Z digest=sha256:af8ceece776f884f492cf5f4b139f540c5d967bd24aa46fcdd8d891d3c0fbfda

Observation bdf259b0-3689-4d5b-baa4-6cf15306f100 · outbound

This paper cites Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.271321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.271321Z digest=sha256:f38538b52e7bea15c6f1874f7dd737dcaccf318ee66349adbd059d6908262779

Observation 17274775-abc9-4f93-85c8-7927546b49ff · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:04.450022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:04.450022Z digest=sha256:479d85d7c8311085376c520b24f95de07e84f8ee66679a8e3791f5077dd0f4c7

Pith citing papers

Observation da5bd0c2-fcbc-43ca-b841-79eaff1e1b73 · inbound

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction cites this paper.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.994077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:fa761dedcbb0771d7e6ac5a1fdf2f7a1a099a68c7ac8c127fef69ba9ac80e383