Pith. sign in

Paper Citation Record · LEDGER

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2506.07081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07081 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:47:14.172217Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:53.938483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:39.127928Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44080555-c607-49e4-a760-a70df34c9de0 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.320572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.320572Z digest=sha256:1501e1d990b4d927b4e900badadf1292506ee1a4e9724f3668abeaeee049226d

Observation 1ca75ee5-98f1-49b8-9bce-9eed04c1a35a · outbound

This paper cites A review of subjective scales measuring the user experience of voice assistants,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A review of subjective scales measuring the user experience of voice assistants,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.964681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.409467Z digest=sha256:45bde0540583cffd1cda4956001d4c86ffcfa1218f7d194069e73276d2ae9722

Observation 54c688bf-1a8f-4aa7-8334-91adfada3c3c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.481844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.481844Z digest=sha256:6f4b7ea4915635ed063ff7affbc19d62d9c60035410be6ea7b04739fc7cb8d3e

Observation 0f0da420-5713-49ff-9605-d978e32a8c85 · outbound

This paper cites OpenAI-gpt-4o,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training OpenAI-gpt-4o,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.794138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.526889Z digest=sha256:879b070bd939c45b9aca59f225a1e7fb3fde7851d5231d07561326aee37b2ad2

Observation 261a5a25-60ab-49de-b987-497f08b38a76 · outbound

This paper cites Improved End- of-Query Detection for Streaming Speech Recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Improved End- of-Query Detection for Streaming Speech Recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.682409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.656767Z digest=sha256:2f0855504a97af6770b7f0da2883f6c07ef90a4e2f6c7e6791a3089e408afeca

Observation be2a5081-a9d9-4e68-9d62-720bfb6fda20 · outbound

This paper cites V oice activity projection: Self-supervised learning of turn-taking events,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training V oice activity projection: Self-supervised learning of turn-taking events,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.539409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.711258Z digest=sha256:f00c693e599b664989dff40b22e9fd9b4b2c821e3313adec42bec537c8a01ba1

Observation 651c22f7-0f2f-4f6f-842d-169ef6f7cf64 · outbound

This paper cites Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.396827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.779580Z digest=sha256:30363ef2fd89fcb49973afe159d82e3b82abe0ce40736ad84d27d4ffc46a3f3a

Observation bf9ee3b0-9255-4331-84b8-4392ca4b0ce8 · outbound

This paper cites Root causes of lost time and user stress in a simple dialog system,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Root causes of lost time and user stress in a simple dialog system,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.274468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.855592Z digest=sha256:978b2baebe2d92a249b7fe56d9c70ce1f8ec190d2c727be5c6817b6add77ed12

Observation 96915e7a-187b-4b5b-b6f9-52620342af30 · outbound

This paper cites A statistical model-based voice activity detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A statistical model-based voice activity detection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.096855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:11.946892Z digest=sha256:67d2cb2c465565017b82731e2087abd0ed4f9a7fb0eed19a6bfce9af062b0d67

Observation 1f4f2787-25e3-4410-b0d1-4d46fba5cc3d · outbound

This paper cites A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.965610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.001979Z digest=sha256:879534eab8acb4b149edd97b58353fa4e640b88855437c09066f2b024a59755b

Observation b96e219c-d423-4ddf-9bd5-d3b65f4c9073 · outbound

This paper cites Temporal modeling using dilated convolution and gating for voice-activity-detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Temporal modeling using dilated convolution and gating for voice-activity-detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.753204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.092752Z digest=sha256:dcf10b15502a843f1bee0a40c7e47e544b4d492d15208d3b0dc2e57479c8770f

Observation 7ad4af9c-0153-444c-8f67-863e23fd20bd · outbound

This paper cites Robust end-of-utterance detection for real-time speech recognition applications,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Robust end-of-utterance detection for real-time speech recognition applications,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.629364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.157555Z digest=sha256:0d1e87e95f556fc91fb980fdaa76324336ae1224c98329289d0e8c091bed2123

Observation f50f5b3e-03d1-4f6b-8d10-1ec217bbbdf2 · outbound

This paper cites Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.503583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.248057Z digest=sha256:ae0bea00189e47d03404f0f67ba20e4d2e50c37b7dfb390a7145bffcfa4d1da9

Observation 5dc9aed4-5c79-4292-830b-050f5772222d · outbound

This paper cites Dynamic speech endpoint detection with regression targets,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Dynamic speech endpoint detection with regression targets,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.308518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.341985Z digest=sha256:05f9b037ece3b53fc5cfbf5d629fe568b8114f151a88ec650c5a53e2e1058f45

Observation b86fa938-e980-45c9-a34f-227472dbd7e5 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SoundStream: An End-to-End Neural Audio Codec,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.141871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.410195Z digest=sha256:5043447b95d8bbd3fc79b40089cf3475382e5d50a68c6807b0c297c9f3735e40

Observation 34bd2dc7-2de0-4b9d-a23b-3ffbca6bbfe4 · outbound

This paper cites High Fidelity Neural Audio Compression,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High Fidelity Neural Audio Compression,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.978334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.520620Z digest=sha256:1c377258a9ad1d21a405cdccf050c30e63382d4b1e325080f6a3419fd1fd573f

Observation 17e1c7d0-bc0f-4cfc-81f8-2d25bb286c4e · outbound

This paper cites Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.830311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.585964Z digest=sha256:d474f8385084a4aa7add7e14a6ec19bb30c0a8b40285c8241b5a204ad963a713

Observation 9d41f697-b6d7-4954-95f6-dc506c121022 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Moshi: a speech-text foundation model for real-time dialogue

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.652698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.652698Z digest=sha256:3587070d503c06c828e0e74c6acb36bf609ae94af60f38813e3b5bbaa7e14591

Observation a363b232-9cb9-456b-9fd4-f460d74c6255 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.616111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.704887Z digest=sha256:5905d73a607eb2e92d9e3b801ce78802e2d6a96c09600eacbf7a66c62cab1be2

Observation 6595c748-2f80-4c2f-b947-401edb609c36 · outbound

This paper cites ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:47:14.407787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.794875Z digest=sha256:5296a3070d1cdef25b467da4ca11d9d3c14d2b7fd700e5f8f7592d83a940da7d

Observation eae1316f-4479-4d94-9707-99e0bbf49b08 · outbound

This paper cites Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.453168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:12.857921Z digest=sha256:218ace3ed361ad5d901077822689ea1558ac000a5995a4ba39a5f146960f045f

Observation 9893ded1-943e-47cd-b054-faab046ad98c · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.941618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.941618Z digest=sha256:36fcd86f202bfedbf184b140066c18f6110bbcb42814d22711c1f765493b3b15

Observation 63e01f50-c2ee-408e-a658-eb4bf3d9617c · outbound

This paper cites Self-supervised speech representation learning: A review,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Self-supervised speech representation learning: A review,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.003105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.003105Z digest=sha256:53e81dca8cf53fd4240d9a9b51c37d23b0d8941ee7b4a016f5998119872922cb

Observation 6a621974-0784-420e-8bea-cabd6897ef01 · outbound

This paper cites HILCodec: High-Fidelity and Lightweight Neural Audio Codec.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training HILCodec: High-Fidelity and Lightweight Neural Audio Codec

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.130003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.130003Z digest=sha256:06ec492cfe8ec97803d7a936880299fd9c84053dfc1f557c3c0249c43b3567da

Observation abbbbc0c-103e-4429-a95f-dbb4fbdbbd95 · outbound

This paper cites End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.289254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.214580Z digest=sha256:7c31f7173be82dc3fd3f693a8f6ec80372697e32a0453bb65ba9e51c9fca3922

Observation d98b2206-b1c5-4486-9f80-464af49ccf38 · outbound

This paper cites Joint endpointing and decoding with end-to-end models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Joint endpointing and decoding with end-to-end models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.150228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.254678Z digest=sha256:6696195da55e5bf5bbdf1fce0ccb24fb3885b120db5adcf1f8eaf7c40b9ecef9

Observation 4344bebb-23e6-4ca9-83c1-9f4ab23c83a8 · outbound

This paper cites Turn-Taking Prediction for Natural Conversational Speech,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking Prediction for Natural Conversational Speech,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.992507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.295077Z digest=sha256:ce6aa91184e1ee2d06bd2db608f2b6ec61bae9f46907e198ffc59575d2510a1d

Observation 1aad363c-8c0b-479f-872f-b23e539ee1a3 · outbound

This paper cites Towards fast and accurate streaming end-to-end ASR,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards fast and accurate streaming end-to-end ASR,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.879748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.378298Z digest=sha256:6cbe22b53f9251bff2ba30b9b80db92c62d0a7a81fabfd4f03c728da2b2df9f4

Observation 30dea1f6-734e-4b63-899a-fcc8d19ed347 · outbound

This paper cites Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.705567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.441465Z digest=sha256:c5c63e8ff073985caaa558e03b2f2219a05c17e01a7cfe444128df72ed30615d

Observation b5881084-3c11-4493-bfc4-749fdef1cbd7 · outbound

This paper cites Two-Pass Endpoint Detection for Speech Recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Two-Pass Endpoint Detection for Speech Recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.631628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.504091Z digest=sha256:bc9e366061b4c11a8f6293ca67590743cd425bfab887ee87d9a645b5a1ff5c81

Observation f50655af-451a-4b4e-9481-990a3529c45f · outbound

This paper cites Towards Accurate and Real-Time End-of-Speech Estimation,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards Accurate and Real-Time End-of-Speech Estimation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.506542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.535657Z digest=sha256:39a810b676b4b7a9146f7d51811621ae2fce3c77f5159aa0bb8af8b168d26ec7

Observation 473084b1-adfe-47ea-87a7-353b44ab619d · outbound

This paper cites Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.264732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.576137Z digest=sha256:bb3a5945ca93a441927f8ce3d632a9da3bfbea477933b6fad729967a34cb8bbe

Observation 7e73b4fd-9911-40ce-8881-1abafd2b4420 · outbound

This paper cites Multilingual turn-taking prediction using voice activity projection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Multilingual turn-taking prediction using voice activity projection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.112289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.643938Z digest=sha256:a489015e23f33694cadcb0d031aec0d29c4f0d21049755e2ac2cfc6a4c5b2672

Observation 89e63db1-7533-4c20-be69-3ed56f734810 · outbound

This paper cites Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.914320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.681171Z digest=sha256:7ba71ae788b67caa08be32e4d4064e8f7575c3a245c0a93a65f220eb524856dd

Observation 45cbefc9-f0af-4ab9-b512-200a9d8bfa12 · outbound

This paper cites Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.745811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.728539Z digest=sha256:0ad3fd1a7c8f487c02b0cc6aa2706739edb646a1f8e5f76817ba3410357e5fab

Observation 9c4f5350-1927-4723-af50-68fd90cd8876 · outbound

This paper cites End-to-end automatic speech recognition integrated with CTC-based voice activity detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end automatic speech recognition integrated with CTC-based voice activity detection,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.537482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.756574Z digest=sha256:0cce61374d2143b527cd2077f80c81a9289db59426085332ed4ece0d052e0fd2

Observation 4debdea7-6324-49e2-aece-de6ddf6788b2 · outbound

This paper cites Text- free prosody-aware generative spoken language modeling,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text- free prosody-aware generative spoken language modeling,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.809131Z digest=sha256:e7f77b31ccd27e59c93a4c6d006131e8b0d904d4066712c6b8672fa4abc0a978

Observation dc82655a-cacc-48b4-b0f3-ee6fd9074f98 · outbound

This paper cites Simple and controllable music generation,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Simple and controllable music generation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.129513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.835391Z digest=sha256:a4d04051314810f393bb4f8eab43cb4892e2614df0b45afea1498ea8005cfab5

Observation 43085501-452b-4b50-97a8-6d7ec417e90f · outbound

This paper cites SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.062955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.880359Z digest=sha256:8464a043a5e20e47e2f9923e02d92f3d7b58ede12a0aaa8cddea48805ab32303

Observation 60401cb3-9b9e-4899-8663-ca7f1a9d073f · outbound

This paper cites Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.921334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.921334Z digest=sha256:a14a781aebbd0972f8d33e8b120ffb38d0bfd140c8980f67fddac2b370f996b2

Observation c425acdf-4ba8-4ec4-9a02-d5a7cb25ca4e · outbound

This paper cites PyTorch: an imperative style, high- performance deep learning library,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training PyTorch: an imperative style, high- performance deep learning library,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.816798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:13.960499Z digest=sha256:eb452cbafffcd1929ea0bfa1f2e7a30579d419818301101546ab3e20d56e1860

Observation fc38f094-28c5-4cad-b465-4167ba175604 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Transformers: State-of-the-art natural language processing,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.679549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:14.010240Z digest=sha256:d69d7ef5f79042e01da1031e0a12d7861ba2d4fa954537112de2bce23ea12a75

Observation c25e00e3-6b18-4177-a50b-80c08dcf091c · outbound

This paper cites Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.564865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:47:14.075639Z digest=sha256:684c184987951ae3e96a3368f29307f43807cf35f5c6d669ac0878666d49010b

Observation 2768c705-0e2b-4bd9-9b30-d63fb9d189e3 · outbound

This paper cites LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.122383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.122383Z digest=sha256:1e89be31490feaf5fc4ece8dd6babc2868672fc3d9d9fb0c4b13b47872dc5e02

Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.172217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.172217Z digest=sha256:dd7c8455725a1b0d10ab52ab11e0ae20b46742c884763a629db462193ee26faf

Pith citing papers

Observation 291052d3-5465-429d-be90-0dba25588d20 · inbound

KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI cites this paper.

KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:53.938483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:53.938483Z digest=sha256:9cafaf4c89d0f11fcc885420aaa7db77aca32996f05e886d7f42673efbb83364

Observation 1f8ee3e4-80fe-4a9d-bb4c-5c66b4c1ee93 · inbound

Endpoint Anticipation for Low-Latency Spoken Dialogue cites this paper.

Endpoint Anticipation for Low-Latency Spoken Dialogue Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.129488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T05:37:17.684884Z digest=sha256:b3151e6f2032dd2b84df8c18938ce9f84701e6fa9f48ff024b1c201b2fee3fe3