Pith. sign in

Paper Citation Record · LEDGER

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

As of 17 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2506.00975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00975 v4

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:42.951304Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T08:53:13.408944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.307813Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c6e1d7f-e66e-450c-b2e6-694dcebd8e65 · outbound

This paper cites URL https://chattts.com/.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction URL https://chattts.com/

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.567763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.567763Z digest=sha256:411571d695d76f23c7d09d7d9e462aa48a07614373ef01937f3c19b484954565

Observation e29e318d-d351-4b0b-b673-5f2dafab736c · outbound

This paper cites L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.619472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.619472Z digest=sha256:1a161269423ba58e058794e600de9206bbd675c7ff3539063b86c9dfb423d822

Observation 6c068f11-9a6e-4a59-b3d1-ebd2798c923b · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.751404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.751404Z digest=sha256:e3a82d56501d9d128b39c7fe7c286c56f3401fce5ace6d5adf25d09e82345274

Observation 895298cc-c608-419f-998e-e946641413b7 · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.860327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.860327Z digest=sha256:e773e869529c0f7b7b755290f1607200c0a9448d1185a5505388cb548d9b96d8

Observation a1636c54-3f43-496b-9d53-5b339f20af2f · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.997075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.997075Z digest=sha256:4a3c82b3e66e407820d67b2c9395c6ea2e3bce9ec5d4a86b3584bfbe68d6ccd9

Observation bf5438fe-1fe1-4d58-925c-5c8a7a0fdcdd · outbound

This paper cites UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:58:43.660736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:33.119159Z digest=sha256:ae8f6a6c9d4468f6fa838ec8fba21234c5d36b26d2322c4bc44c6ddcca4b1f46

Observation be36cb6e-4e8f-4c7f-a43f-6d65697acb32 · outbound

This paper cites Wav2vec 2.0: Learning the structure of speech from raw audio.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Wav2vec 2.0: Learning the structure of speech from raw audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.224038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.224038Z digest=sha256:243d3a52db35ffacbe44c094b058c9d99eaa036ee753af1330eac81ac802058e

Observation 27ea4246-ec6f-4fc2-9b03-b0f163153d80 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiolm: A language modeling approach to audio generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.386322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.386322Z digest=sha256:4559cef91564d218d25ff4ff65bc209b2562847a589563c78858aab9d24e2e18

Observation 7bb04f5b-c50f-4ee3-982c-7d70c3d03ece · outbound

This paper cites D., J \' u nior, A.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction D., J \' u nior, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:56.927864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:33.522304Z digest=sha256:05ba9cd3307d9028ab390df69d74837e533271092a3f2582c8c413fe2ef64a62

Observation 9e0bea6d-ba96-4772-ad3b-b985ffe392f6 · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:56.691490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:33.640729Z digest=sha256:d9b0e24557e435d4562e2561615bf8a85c9450d08a0ddf05bfbd0b19b933bf30

Observation a3105a6d-23ac-44bc-adf1-d7dbaa5fd501 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.743754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.743754Z digest=sha256:c2fc1560a49f64bd9551477592102152c52111552825733b4004f203ebe3a991

Observation 78a49792-6902-455d-b0a8-d4bde5038fe8 · outbound

This paper cites Qwen2-Audio Technical Report.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Qwen2-Audio Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.890548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.890548Z digest=sha256:93956a2dbb9d8d78045ebc309e9d8ef44939b99e28bbb58b04548bb3f488c5c9

Observation a5a904c0-bc0b-4e51-b1b9-2c30286229e6 · outbound

This paper cites Fisher english training speech part 1 transcripts.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Fisher english training speech part 1 transcripts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:56.470785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.008369Z digest=sha256:7dab14b21b7b8138bf9698ef4dfc6ad6fd4fe162425ab7ea6d92e37c2ebd81b5

Observation 22a12f79-7fe4-49e1-8e7f-61dfd7ec94ef · outbound

This paper cites Simple and controllable music generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Simple and controllable music generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:56.194214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.130601Z digest=sha256:a2e6f9bd640387f42ac1c8dd8ae17ffce4474f8b1cd85cca15d639fb0a869e4e

Observation 312dfe16-cff2-4271-8325-6e2db28426ef · outbound

This paper cites Y., Ermon, S., Rudra, A., and R \' e , C.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Y., Ermon, S., Rudra, A., and R \' e , C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.926907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.241000Z digest=sha256:1c938dd0b64d10cab53620b9e5abbb18aac2208f7d478f8669e9d3ab26f4fa3f

Observation 7d474f38-9b34-4734-a226-eb1c4ca373f0 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Moshi: a speech-text foundation model for real-time dialogue

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.619710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.396487Z digest=sha256:cd537bc56537eb74afc7ce5eb307c1d6bb60ac9ee3adf1c633b1a773e2d18769

Observation 8d72e998-71f2-475c-9171-1452dc0e2651 · outbound

This paper cites Pengi: An audio language model for audio tasks.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Pengi: An audio language model for audio tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.424225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.505924Z digest=sha256:3dedc9ef9e05bc075bb556456dbcde21efd4089dd9237b785510a0ffa9af533d

Observation 8035faf3-7c08-4760-b83f-b6f966129199 · outbound

This paper cites The Llama 3 Herd of Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:34.644720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:34.644720Z digest=sha256:19afa6b85346dd59a3ba8938ab4968bbfe1a7a4bb44a9bd0085fbe83dcd2c95b

Observation bc2143a5-894e-4cd2-8ea6-84f5f8324f3d · outbound

This paper cites A., and Wang, H.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., and Wang, H

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.196143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.781530Z digest=sha256:522c2570bf5a58d16404f4d46927d25bf3ce8b75b25523c32dfe87848f8f5f97

Observation 92319e28-3802-475c-9737-cb73be4d4fd7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Taming transformers for high-resolution image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.968188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.905358Z digest=sha256:d53fd22b4f210e57a694d08e648015128dd489d37b0571a61a435992f3cee062

Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:35.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:35.040733Z digest=sha256:6fe83ff3f920bf335ceeef51179ea122bd5a3a0a9ac20a2dd5834c2cdebda53e

Observation cbe33887-8a6b-4137-8a82-948411c9230e · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiochatllama: Towards general-purpose speech abilities for llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.677786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.199376Z digest=sha256:bacb867b0728ac54ad05a33442e1983f0e0f1ad9829925c7081fe561882bef78

Observation 3827e764-6204-4997-b37d-b787117fdf61 · outbound

This paper cites Vita: Towards open-source interactive omni multimodal llm, 2024.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Vita: Towards open-source interactive omni multimodal llm, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.488112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.296527Z digest=sha256:770c737425dbf5506524f1fa03fbc48aded5f5053e4c49f2aa5ce3529b58c21c

Observation 3af4459a-7331-4898-9375-ffa2d8b34895 · outbound

This paper cites Funasr: A fundamental end-to-end speech recognition toolkit.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Funasr: A fundamental end-to-end speech recognition toolkit

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.283084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.436272Z digest=sha256:bcca653e977e6a70aca70fafb0b0b6a6c19df3a1ae29091b5edc8f0f70ecd59a

Observation ac321634-5ca9-484a-8c90-13f474f18b1f · outbound

This paper cites A., Gat, I., Conneau, A., Kreuk, F., Copet, J., D \' e fossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Gat, I., Conneau, A., Kreuk, F., Copet, J., D \' e fossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.961536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.569886Z digest=sha256:3702c745955d28a432b07cf4eb24a2b6922135beaddb54f698a64dfec669725d

Observation 551d551f-dee4-47fd-8ee6-60bd884b59d8 · outbound

This paper cites Kenlm: Faster and smaller language model queries.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Kenlm: Faster and smaller language model queries

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.726981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.700587Z digest=sha256:1073bee2cdafdcafbedb61f551f372457e7b2a5cf75675afbf8cdaed26686ca4

Observation 8dc67ce7-02a3-4c47-882b-a35b76ec12f8 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.570099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.899089Z digest=sha256:4b631f41decd32d0f7d2e6ec05e444ae21723a5cad02044de56c9c47b08ae059

Observation 70df1984-f0c7-4d8c-8fcf-ba11ed3134da · outbound

This paper cites Mistral 7B.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:36.013224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:36.013224Z digest=sha256:3ae30d08b321f13dfd089fcf0a149c27d301b81d06b6f61a4088c44ee7f73c42

Observation 3edb39cc-aa46-4f93-b2bc-0f6f5ec4b43f · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:36.096538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:36.096538Z digest=sha256:4e8915b1077421afc12c9b673a6622fe870acd5a253b6ba2ffb589d1fb6ebda6

Observation bc7481c0-d0d8-401b-a7cb-ccaa79b20de2 · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speak, read and prompt: High-fidelity text-to-speech with minimal supervision

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.286789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.223271Z digest=sha256:a0c62aed09e934e18799ebab2de23489f36189b58fb4106dafc337b61c89b02e

Observation e6a44f7e-c36a-485a-8fcc-6b0313bede3d · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.057115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.333973Z digest=sha256:a1f215fe7163e481409fae8cb5a4e78b2fe9e4110ce973a8e0a93b10f40893aa

Observation 4acb6830-e1a4-4ae2-9dd2-108b2a38de82 · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Diffwave: A versatile diffusion model for audio synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.777998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.431685Z digest=sha256:34a59849e8cc37a09653387f2f6ba59bbc37012b624b954ab9a7d458a3caf2d4

Observation 8aef1af1-6c15-46ed-aec2-25318c405e47 · outbound

This paper cites Audiogen: Textually guided audio generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiogen: Textually guided audio generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.568527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.542580Z digest=sha256:8b243e21be27c578ce081fb903b9e83883b63881b257c2f60b1ed4f140d2ae43

Observation 99c49b36-5fe5-49e4-9401-1637cbae224d · outbound

This paper cites Voicebox: Text-guided multilingual universal speech generation at scale.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Voicebox: Text-guided multilingual universal speech generation at scale

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.376290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.660333Z digest=sha256:c89171fe3fcea02b2ec6209e5e58a64eb021b39cbed37749cc163c244ceb68f4

Observation c18203e5-cdca-4886-98ff-5af9c24c8ace · outbound

This paper cites Autoregressive image generation using residual quantization.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Autoregressive image generation using residual quantization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.151875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.826730Z digest=sha256:156688e03b08809eb10a529ba4fa32ecfc4f05434678d1f46305ba944fd86b15

Observation 9cada8b0-4ef0-4bcf-917b-c8d5e96ebc34 · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:51.953854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.965426Z digest=sha256:e239c2d8ef777279419b64ad866f933a67408ae20643e27d645b5b85a9e87e8a

Observation 51370c96-703f-4c0e-a66f-2859258b346f · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:51.714687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.080037Z digest=sha256:7ae4e2b31549d6b593bec0b5da17c8d7c4f2fbbf9b9e3302445511d24dc812fc

Observation 7cd156e9-0f03-46af-80cd-2e2aa83ec3e0 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Autoregressive Image Generation without Vector Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.209661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.209661Z digest=sha256:62c3a92ffbadf06398b7f2738b294ff1e6e4b11768e48ae765cd0727dceb3b6d

Observation 5b84a7b0-9973-4d37-b53d-67fb7fac0404 · outbound

This paper cites Evolutionary-scale prediction of atomic level protein structure with a language model.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Evolutionary-scale prediction of atomic level protein structure with a language model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.351317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.351317Z digest=sha256:3603c4e17e29d526a12dfbcd5cd9aea2ffb99f750a012025fb02ee93736cbea1

Observation 577737c9-fc2d-49ef-a029-31c956a9fa82 · outbound

This paper cites P., Wang, W., and Plumbley, M.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction P., Wang, W., and Plumbley, M

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.422593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.486134Z digest=sha256:e616fbf3e5b18248ce19114c6174c67fba69eb308b25f142944760aa03454c74

Observation fe9c039f-14ea-439d-a24e-e7beb456a9aa · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:51.155503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.574388Z digest=sha256:024dc3c8e20ebc1aa915654dce3dd16a5a63567039c4bea9f18d795991294426

Observation d6e50ae6-0bd2-4ffe-89d5-8c9e59471355 · outbound

This paper cites Language Model Can Listen While Speaking.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Language Model Can Listen While Speaking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.678428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.678428Z digest=sha256:3064e80a724911ea6004397b4006665a7dce5d300b18a510b27d064389b2bd99

Observation cb002973-bc68-4d7f-9ed1-0634d35b79a6 · outbound

This paper cites R., Subramanian, S., Mohr, B.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction R., Subramanian, S., Mohr, B

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.887528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.804722Z digest=sha256:c2a0a3d355a6534cbab56ca9be37595d6791afafae4863ba351604c2b943ef2e

Observation 75dea8af-703c-4fbc-ad56-db0becf0e6c3 · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.901004Z digest=sha256:9d3539e55fc9c89ee957026a708ee1e2b9119a9583c784a0fcd33e19dd8e108a

Observation d6f5d090-d7d9-4f53-aac1-56f46012cc2c · outbound

This paper cites J., and Ramanovich, M.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction J., and Ramanovich, M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.643978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.970581Z digest=sha256:4ad0f5edd752c23112a41bcac16fd7e182f654c441639091d5d7700709eddd11

Observation b78aef10-d1ba-489a-aa7c-b115592f2056 · outbound

This paper cites A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., and Dupoux, E.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., and Dupoux, E

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.461537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.153472Z digest=sha256:33ea0ca97d943c5c1c5243f95a0a9bf408d1c3b715897157c02268cae351acc5

Observation cc02d0d4-486c-4a1a-8ae7-bc2233c0cdcd · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spirit LM: Interleaved Spoken and Written Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:38.270437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:38.270437Z digest=sha256:255878b9a1341d7d1e24876ef562220a56287491fe4611500e75d39c2ad48ef3

Observation a8d38a4a-ce27-44c7-8e0b-09a04e20b094 · outbound

This paper cites GPT-4 Technical Report.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:38.403257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:38.403257Z digest=sha256:051cc0323381211fb49688e16569bc0fc7131f71c71003edd2bc195b1f4a6cbb

Observation df8a2e81-3e6c-49fc-9f93-86634a4510a2 · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:50.225321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.537332Z digest=sha256:f9913da96ed72826203e3962f33c69d914117eae806ce4b9cd53b049b639dade

Observation 27299d89-e9bc-4255-8f4d-c4c55c3362ed · outbound

This paper cites S., Constant, N., Raffel, C., and Callison - Burch, C.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction S., Constant, N., Raffel, C., and Callison - Burch, C

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.986362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.665681Z digest=sha256:385aae8ea5130b2a8f5c4ef1c83712ddc2aa42d4bcc7eb5988a994ff1dd5d43f

Observation f1cd5d8c-5ea6-4c03-856d-6bec1acbf740 · outbound

This paper cites Efficiently scaling transformer inference.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Efficiently scaling transformer inference

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.691407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.778622Z digest=sha256:10ae5868e2b6ceb92bac638a14b70ee27341e9813c174f1cd07325f1813e99ec

Observation 42664b0f-f8b7-4927-9566-7a989af13eb5 · outbound

This paper cites Efficiently scaling transformer inference.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Efficiently scaling transformer inference

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.426335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.895582Z digest=sha256:a405410403fbd19f50b7ff04ec4aa3eed9fca5c29fe407547c69fa544721a12a

Observation b1d2de32-f4f6-489b-80c9-03afcfa5fba9 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.176260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.038734Z digest=sha256:9b280e47f451c122d0851aefbb7e1b975b5289ddbb07db2ff2b330fda8169d12

Observation 5bca6c3d-f9c0-4063-a333-af37c71766f6 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.949112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.145468Z digest=sha256:cb6ae88bcde67b7fffd5d321597d48bfd30e38d65485b106c209c7bed96ec6c3

Observation c0a44f52-f6db-4094-9d7c-9c56bafaca96 · outbound

This paper cites The candor corpus: Insights from a large multimodal dataset of naturalistic conversation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction The candor corpus: Insights from a large multimodal dataset of naturalistic conversation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.665985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.261569Z digest=sha256:e156be4cce575a939d8d0dd4ea0bed1bcd9fe1f8e6775e1912e6ee23bbd7eb34

Observation 7f46aba9-72bc-4c48-9fe1-a15b3e8476f3 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:39.378115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:39.378115Z digest=sha256:717837f86a90cd7b4d2ae4eed08dfea861816ad092dba795bad1a0ab7de18137

Observation 251546c9-78fa-4bdc-8f80-5870134b192c · outbound

This paper cites A., Bekas, C., and Lee, A.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Bekas, C., and Lee, A

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.434728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.514408Z digest=sha256:f1f3110111f3f0273ed50e8ecfff29c31a65368f1ab654d736efbc91301d3174

Observation 4be0a642-827c-40a8-89a8-0ade6f0c6960 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.208784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.602846Z digest=sha256:fa9b4bd2eeefbe42a6ab4c088201fdce3549d793c0e3df5f06cd72689e647dfe

Observation 4458dc50-c08b-4f9c-9dd6-67226ad76cde · outbound

This paper cites Graphaf: a flow-based autoregressive model for molecular graph generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Graphaf: a flow-based autoregressive model for molecular graph generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.985599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.699015Z digest=sha256:6085c150504eff858e74701dee98b63214aa28ecc6963148e7663a90bb1e70ac

Observation 76212d70-838c-4c6e-88c9-e8652cd920d8 · outbound

This paper cites Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.748924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.858181Z digest=sha256:0d88be6549ea94ec4ceb983a9620b15cdfbadd08901b7836feecbcc29b6a1fcd

Observation 0a36bdfc-2d8a-4785-9013-a3753838569f · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:39.979331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:39.979331Z digest=sha256:7a3767239b834861aa0743562f21802ba1f308e49a374f45f4a31396f6cc4cd5

Observation 2f64fc8b-29c9-40b4-b397-55cc1f6fb24d · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction PandaGPT: One Model To Instruction-Follow Them All

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.100604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.100604Z digest=sha256:46717bbc2d5acf261fb9d94cddb8f2377d90feca65a0b9432c86720ea919fb8c

Observation c0eecd77-e1a7-4f59-aac7-d6522552d8cb · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:47.513859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.217367Z digest=sha256:0e29ae26f436dd351bbf969a3790a9119ca3b7e3c629f3b9652812c860a8b55e

Observation 53eefead-cdb1-4489-aac4-d01aed5c75f5 · outbound

This paper cites SALMONN: towards generic hearing abilities for large language models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction SALMONN: towards generic hearing abilities for large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.257441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.305666Z digest=sha256:192e768dca0272e757771f893e16099fa78d9e45a191654dfcaf1a45dbe0fe9c

Observation ef1f5c20-0cd7-4bb2-a8a8-2266bcff8bb3 · outbound

This paper cites C., Ture, F., and Lin, J.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction C., Ture, F., and Lin, J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.008096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.406006Z digest=sha256:dc2e6f18d5d195d3d82ff8fe370dbbadec7f6401e1b97d205e5d97864c254759

Observation ead31476-df3c-467a-b001-a59e2152fb08 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.519186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.519186Z digest=sha256:410be522f7100693e683bf8c12eb47eb873476efe16e2b570adbbe9a889bca56

Observation a3628434-6402-4c47-b395-c74b301bb76f · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Gemma: Open Models Based on Gemini Research and Technology

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.614181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.614181Z digest=sha256:c832bdf956923b6c32ea23a2c0f555aca710ed567edc3919a0f5352fb4ac5c3d

Observation 9b29c5b4-6da7-4c8d-952b-b502da87403a · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.740546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.740546Z digest=sha256:587b7811df85998dd23b9f2951730b4c565d8ff82e361c79e29a21fc7364f50a

Observation c6014061-9c6e-4f7c-8312-53d06888a81b · outbound

This paper cites Neural discrete representation learning.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Neural discrete representation learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.790852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.872224Z digest=sha256:0454e899c197702d3f59a9670409c3603a63c37fdab59cdfe3265ce39dc7bf9b

Observation 66c2f283-ffc8-4f9b-9b40-b25d6adb6bc8 · outbound

This paper cites Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.965073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.965073Z digest=sha256:c1230faba39593b8b8e26929c839c00afcc4155df2cef4a419f73a674de2a646

Observation 4870119f-084f-4438-ab2b-8d42b446748a · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.093509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.093509Z digest=sha256:a318bb3aedc2d2096a8aa7b78448552a719a063ef2a2dad9197f2d3b34e450b3

Observation 30b2b3e3-7e84-42ff-a520-9b7525c41262 · outbound

This paper cites a ckstr \.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction a ckstr \

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.543674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.201189Z digest=sha256:40cacf0e3446397e23db256c6ac1fa436e4c39fe0919dd28b4a0e091875a36ef

Observation feb87d25-b143-43c8-b5d0-e1124abada8c · outbound

This paper cites SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.346118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.346118Z digest=sha256:747fc761c231f2320940b6527d797cf5d22e5b2f151cd47b6411717731ee4d22

Observation 7218aea1-0087-49b5-831f-f71b3e529ccc · outbound

This paper cites Next-gpt: Any-to-any multimodal LLM.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Next-gpt: Any-to-any multimodal LLM

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.249196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.460790Z digest=sha256:6aaaee512ff068f777b9350a418ad6e9d7ad633a4bfc22eec8079478f5d3a12c

Observation 682e39be-d364-4755-8a56-0b5f53469a1f · outbound

This paper cites J., Wang, W., Lin, K.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction J., Wang, W., Lin, K

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.929927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.609319Z digest=sha256:d30d41634349e92d8bc8c2a3b021608fa559c5078f947efebab600f02a68b101

Observation b034d105-1558-44c9-8ac2-b0fa4e255f43 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.742682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.742682Z digest=sha256:dc4662671ebff0c81ff6506399d79e1631d5213093211971fd3828b26126a933

Observation 5c7c0083-2de3-4e05-9bd1-c91a2c79b57a · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Diffsound: Discrete diffusion model for text-to-sound generation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.640110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.844624Z digest=sha256:70caa9f4dd3e9440150477a739e7000cdd4c0478817d300c4c725fa9afa73357

Observation e3e54503-2c5e-4ae0-946b-8f85d0a409e8 · outbound

This paper cites Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.384003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.002641Z digest=sha256:28d7a8de1d84ca3547f10951f802f6492dfd4d0a7fb8d726034278732cd7487e

Observation 3596fa4e-db4f-4031-b596-9e8ae20c528a · outbound

This paper cites L., and Leskovec, J.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction L., and Leskovec, J

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.100642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.157901Z digest=sha256:964846ea8eba1db73f15aa4064e92cc80accb43c0744a51da1e82e5c52fec141

Observation 4e621224-e9df-44e0-abb1-1c2d5ea4649a · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Soundstream: An end-to-end neural audio codec

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.854884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.301302Z digest=sha256:f8ef7d092e9457b1871f902ee45c152b9cb967c70ff39d32b6d568ae12757b3c

Observation 0ce57509-82bb-4136-938e-900155ee3eae · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.569687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.412141Z digest=sha256:ebf1e3baa848d5b134953385ada60166ab82c2b671ad80ec0a9b4a778cd3d9bb

Observation f1ef71db-56d8-4cd7-bc49-777c03c8b47a · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.528594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:42.528594Z digest=sha256:3fbcc20513b5f26d64488940bb5de3448e5286e7c82cdb9f41d12ac6f4512ae2

Observation 4a439b41-1ec4-43d5-8231-6bb96448f359 · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.289109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.650407Z digest=sha256:250d60f2b7e6a6fc029ea39a3681c4e7324db0cb9632b1f038e3402fc4940a2d

Observation 997ba1d2-2a75-48b7-84e9-318ae13e46b4 · outbound

This paper cites Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:43.964694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.812383Z digest=sha256:9146911508058a881ef5ca98a3d4218f456cabc7a246e3d5cea7b65bca276d0d

Observation 69a63b19-1c08-4ca1-928d-38e2304b0bab · outbound

This paper cites write newline.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.951304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:42.951304Z digest=sha256:9439b7384ecbd171d4f2055055a0fd334db0600e47d8b62c74b020cceea4cb04

Pith citing papers

Observation 0f6d85ff-3033-4893-b0eb-51c374cca2bf · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.310003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T21:22:38.888528Z digest=sha256:645eb7f6e7f3751b4476fe650107d2534908bef10bd33370f8eb3d79cb0863ff

Observation 05a346ee-4bdb-427e-b5f7-2ae849d7bb49 · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T08:53:13.408944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:53:13.408944Z digest=sha256:d7c8131a5dc806e807f13df84f7b2ee6ded844ed78f9e0086719379990f33f64