Pith. sign in

Paper Citation Record · LEDGER

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2506.07081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07081 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:47:14.172217Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T05:37:17.684884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:39.127928Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44080555-c607-49e4-a760-a70df34c9de0 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.320572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.320572Z digest=sha256:89fcaca3048180a109035ada982c176f4bdd14acf278bd4f01598b0c36669217

Observation 1ca75ee5-98f1-49b8-9bce-9eed04c1a35a · outbound

This paper cites A review of subjective scales measuring the user experience of voice assistants,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A review of subjective scales measuring the user experience of voice assistants,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.964681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.409467Z digest=sha256:8d3a01825a7acd54d888e7f0fdadd5504d2cc22a7c10a59d8e0f39b7b766b15e

Observation 54c688bf-1a8f-4aa7-8334-91adfada3c3c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.481844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.481844Z digest=sha256:3f14e35aa5becaf87f02049a9cb49bf7a9624c6c5848b6577513cc46a4cb761b

Observation 0f0da420-5713-49ff-9605-d978e32a8c85 · outbound

This paper cites OpenAI-gpt-4o,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training OpenAI-gpt-4o,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.794138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.526889Z digest=sha256:65108975d86b165f1a9c22e5110578b4435200975f56388e4ad4130a351a3014

Observation 261a5a25-60ab-49de-b987-497f08b38a76 · outbound

This paper cites Improved End- of-Query Detection for Streaming Speech Recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Improved End- of-Query Detection for Streaming Speech Recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.682409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.656767Z digest=sha256:e3f12cb876d86aa3ee70ae392cf1ed40148f8e94f6efc3b22e6cae18b9b5b0d3

Observation be2a5081-a9d9-4e68-9d62-720bfb6fda20 · outbound

This paper cites V oice activity projection: Self-supervised learning of turn-taking events,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training V oice activity projection: Self-supervised learning of turn-taking events,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.539409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.711258Z digest=sha256:0134d98e3dd7dd031b074c3969cce5ca4c0754d07243d7583a44576ad0ab8a2b

Observation 651c22f7-0f2f-4f6f-842d-169ef6f7cf64 · outbound

This paper cites Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.396827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.779580Z digest=sha256:7580c734fa448725c448e0f8695edc43039b0a1c913a11e86577afaeb0224b9f

Observation bf9ee3b0-9255-4331-84b8-4392ca4b0ce8 · outbound

This paper cites Root causes of lost time and user stress in a simple dialog system,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Root causes of lost time and user stress in a simple dialog system,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.274468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.855592Z digest=sha256:ed95ea435e95f11ccd1427c2fe668c5d231f888633535d8041883b30556ff1f4

Observation 96915e7a-187b-4b5b-b6f9-52620342af30 · outbound

This paper cites A statistical model-based voice activity detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A statistical model-based voice activity detection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.096855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:11.946892Z digest=sha256:f0453b4560bef6c14ea1283a6300af10e3db34ad86682b8354809fcc9e909909

Observation 1f4f2787-25e3-4410-b0d1-4d46fba5cc3d · outbound

This paper cites A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.965610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.001979Z digest=sha256:f06d1d829980114fa36a4ace9b43e8a18bd2663e10d49a756ab3e58a4d6b3bda

Observation b96e219c-d423-4ddf-9bd5-d3b65f4c9073 · outbound

This paper cites Temporal modeling using dilated convolution and gating for voice-activity-detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Temporal modeling using dilated convolution and gating for voice-activity-detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.753204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.092752Z digest=sha256:bcd97afa25d2bc872bdd5ffbf714b4ff81f81dd963a0a8ce1fe8a73d27c90d3c

Observation 7ad4af9c-0153-444c-8f67-863e23fd20bd · outbound

This paper cites Robust end-of-utterance detection for real-time speech recognition applications,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Robust end-of-utterance detection for real-time speech recognition applications,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.629364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.157555Z digest=sha256:15ae943d5e802ccf5c44261ffc87a1709e9470088e3c946d3f6ec46a5a19b546

Observation f50f5b3e-03d1-4f6b-8d10-1ec217bbbdf2 · outbound

This paper cites Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.503583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.248057Z digest=sha256:ad6b38b3e67db1e305e2bbe3505383def60f66e9b5af5f0ac3bc6012a040475b

Observation 5dc9aed4-5c79-4292-830b-050f5772222d · outbound

This paper cites Dynamic speech endpoint detection with regression targets,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Dynamic speech endpoint detection with regression targets,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.308518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.341985Z digest=sha256:1fb55f1ed5c787b8f2c731f014fec64e2486a88e9025983ab5f5505f38a51b41

Observation b86fa938-e980-45c9-a34f-227472dbd7e5 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SoundStream: An End-to-End Neural Audio Codec,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.141871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.410195Z digest=sha256:aa09292be6e0967f8e8947ebb11b4c9b63b78e5a108f9714e54f7ea6de7c03c2

Observation 34bd2dc7-2de0-4b9d-a23b-3ffbca6bbfe4 · outbound

This paper cites High Fidelity Neural Audio Compression,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High Fidelity Neural Audio Compression,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.978334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.520620Z digest=sha256:f66847bcd8b310140f4f9c54ea4fec9daa6976c58f705217066cd3260d0c57f1

Observation 17e1c7d0-bc0f-4cfc-81f8-2d25bb286c4e · outbound

This paper cites Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.830311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.585964Z digest=sha256:99d90375a9bd4f9cc9d11ee51681c2510e88ea5c426e7fcda69137bd580a205a

Observation 9d41f697-b6d7-4954-95f6-dc506c121022 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Moshi: a speech-text foundation model for real-time dialogue

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.652698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.652698Z digest=sha256:7ed0344a77bc9733e1c1692585eea2cf8e5f74e62f1778ae4a62c1693ae53eb7

Observation a363b232-9cb9-456b-9fd4-f460d74c6255 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.616111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.704887Z digest=sha256:b85603d3f31972f80dcbbeab190a355a52b7f545295865584427421255fcdffc

Observation 6595c748-2f80-4c2f-b947-401edb609c36 · outbound

This paper cites ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:47:14.407787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.794875Z digest=sha256:6db37eff2cdbb6953df504e071ad15f512fa434d991c128d4be5b49f20f286e5

Observation eae1316f-4479-4d94-9707-99e0bbf49b08 · outbound

This paper cites Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.453168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:12.857921Z digest=sha256:0126664fb809fd61d48296c5175cc2b9f9e0bdc2c09603a13aa13e4f89cd2237

Observation 9893ded1-943e-47cd-b054-faab046ad98c · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.941618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.941618Z digest=sha256:7877d7484d6bf3d857449a7b931c0b6177daed75dc9245a30decdccaba7f16b8

Observation 63e01f50-c2ee-408e-a658-eb4bf3d9617c · outbound

This paper cites Self-supervised speech representation learning: A review,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Self-supervised speech representation learning: A review,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.003105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.003105Z digest=sha256:75ce98bad0fe558a051bb6f00520a203107dd44c9826c4ca77b7721a292d6c91

Observation 6a621974-0784-420e-8bea-cabd6897ef01 · outbound

This paper cites HILCodec: High-Fidelity and Lightweight Neural Audio Codec.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training HILCodec: High-Fidelity and Lightweight Neural Audio Codec

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.130003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.130003Z digest=sha256:509a3c28c6d1bb58076041b3ce16328ad694708982f44200a53841fa1dfec597

Observation abbbbc0c-103e-4429-a95f-dbb4fbdbbd95 · outbound

This paper cites End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.289254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.214580Z digest=sha256:38f69920c44070b2f0e738ef72bc181eca7f42af3287c4f75c83726384c1df12

Observation d98b2206-b1c5-4486-9f80-464af49ccf38 · outbound

This paper cites Joint endpointing and decoding with end-to-end models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Joint endpointing and decoding with end-to-end models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.150228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.254678Z digest=sha256:104a93a901cad40302b56f2186df8c9ae0b3515a02ddd2357e40d7d685fe1974

Observation 4344bebb-23e6-4ca9-83c1-9f4ab23c83a8 · outbound

This paper cites Turn-Taking Prediction for Natural Conversational Speech,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking Prediction for Natural Conversational Speech,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.992507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.295077Z digest=sha256:5da02137d39394e31fa01d1412726cf69c17772829aadcbfd9f306a93e82bb79

Observation 1aad363c-8c0b-479f-872f-b23e539ee1a3 · outbound

This paper cites Towards fast and accurate streaming end-to-end ASR,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards fast and accurate streaming end-to-end ASR,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.879748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.378298Z digest=sha256:1ed7ba47e1bd98e5adc4fb1bc7ffc353c3730c576ec9817fe486b12b13c8243f

Observation 30dea1f6-734e-4b63-899a-fcc8d19ed347 · outbound

This paper cites Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.705567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.441465Z digest=sha256:6b8027f9850739aa34ad5e8d162f37349a4ded65bdb09a426c1bef1d7709d7a8

Observation b5881084-3c11-4493-bfc4-749fdef1cbd7 · outbound

This paper cites Two-Pass Endpoint Detection for Speech Recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Two-Pass Endpoint Detection for Speech Recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.631628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.504091Z digest=sha256:92417a5f93b23ba42071dcce96034977d80e1f5bbd1ec42e03cb1f32a568075e

Observation f50655af-451a-4b4e-9481-990a3529c45f · outbound

This paper cites Towards Accurate and Real-Time End-of-Speech Estimation,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards Accurate and Real-Time End-of-Speech Estimation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.506542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.535657Z digest=sha256:1a243209c094a612b419cb5fce2f9d3bd2e25d5f0aeae65816de4b3503c3b35d

Observation 473084b1-adfe-47ea-87a7-353b44ab619d · outbound

This paper cites Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.264732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.576137Z digest=sha256:9a51a4e0db79663ccb450fda21476fce4965d11946392700a3cb8d817971dca3

Observation 7e73b4fd-9911-40ce-8881-1abafd2b4420 · outbound

This paper cites Multilingual turn-taking prediction using voice activity projection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Multilingual turn-taking prediction using voice activity projection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.112289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.643938Z digest=sha256:e26aab9e80ece2cd91280fc87d6ee62e1eaa40f117b52d043159f5ef96d20a67

Observation 89e63db1-7533-4c20-be69-3ed56f734810 · outbound

This paper cites Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.914320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.681171Z digest=sha256:1dca41f8993a48c3e820e24143b8fec7f86e6fa82512db50e225d81b4f014fa1

Observation 45cbefc9-f0af-4ab9-b512-200a9d8bfa12 · outbound

This paper cites Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.745811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.728539Z digest=sha256:820266f8913717b9bc79e4c59b9d389f7d2af0dc5e78809e242f697d37a795f0

Observation 9c4f5350-1927-4723-af50-68fd90cd8876 · outbound

This paper cites End-to-end automatic speech recognition integrated with CTC-based voice activity detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end automatic speech recognition integrated with CTC-based voice activity detection,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.537482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.756574Z digest=sha256:6e80ad4a123577574b966589c8fd75743727e9591e066dced7d9af0fd4d5de57

Observation 4debdea7-6324-49e2-aece-de6ddf6788b2 · outbound

This paper cites Text- free prosody-aware generative spoken language modeling,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text- free prosody-aware generative spoken language modeling,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.809131Z digest=sha256:153be21e5a82e02e1d84f4b367f06a2d017e140854a6b799b0068b4818304317

Observation dc82655a-cacc-48b4-b0f3-ee6fd9074f98 · outbound

This paper cites Simple and controllable music generation,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Simple and controllable music generation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.129513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.835391Z digest=sha256:e62c7bc5ef6084ceae4b21b0f65ff626816e5bec339f637053de07d8e30383c1

Observation 43085501-452b-4b50-97a8-6d7ec417e90f · outbound

This paper cites SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.062955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.880359Z digest=sha256:43d3a4206298902c99569b640f1939334d5c7a9030c9d53362a991b9b23a5503

Observation 60401cb3-9b9e-4899-8663-ca7f1a9d073f · outbound

This paper cites Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.921334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.921334Z digest=sha256:d016f32bb4f680908db1225bd7c3a9c9ef8d84642a51183c813b061139aba0c3

Observation c425acdf-4ba8-4ec4-9a02-d5a7cb25ca4e · outbound

This paper cites PyTorch: an imperative style, high- performance deep learning library,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training PyTorch: an imperative style, high- performance deep learning library,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.816798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:13.960499Z digest=sha256:cbec3c9101cb0a84cdf97c02d1e57f994a4940e94ff3ceedb0afcfbd74373ded

Observation fc38f094-28c5-4cad-b465-4167ba175604 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Transformers: State-of-the-art natural language processing,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.679549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:14.010240Z digest=sha256:ac3ee2c27403a56568c2e5ee53f0733f6368eb9523e739b63d3f74d0ad8dabfc

Observation c25e00e3-6b18-4177-a50b-80c08dcf091c · outbound

This paper cites Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.564865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:47:14.075639Z digest=sha256:4ec20f3dd403f14396c566c7c2a3aed7fc4ee2e4f3a3b97a78a34a4f4bfc9d0d

Observation 2768c705-0e2b-4bd9-9b30-d63fb9d189e3 · outbound

This paper cites LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.122383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.122383Z digest=sha256:599099b92370f36d72dbd7feb4faad4c1a3b9ccd5e0af24056b1e319d0b566c3

Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.172217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.172217Z digest=sha256:b108bdff20f484b2f636ed327a374b7f64d4ec5644fcf48f2e15fbd7f543905c

Pith citing papers

Observation 1f8ee3e4-80fe-4a9d-bb4c-5c66b4c1ee93 · inbound

Endpoint Anticipation for Low-Latency Spoken Dialogue cites this paper.

Endpoint Anticipation for Low-Latency Spoken Dialogue Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.129488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T05:37:17.684884Z digest=sha256:048f1bde51791b6f7a523d85fee08d87040d214aeb7f8753262a0286967e44a1