Pith. sign in

Paper Citation Record · LEDGER

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

As of 10 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2506.00975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00975 v4

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:42.951304Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T08:53:13.408944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.307813Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c6e1d7f-e66e-450c-b2e6-694dcebd8e65 · outbound

This paper cites URL https://chattts.com/.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction URL https://chattts.com/

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.567763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.567763Z digest=sha256:67784c3491a4903bcf52057ecd9e7bebaecbbdc4f7004fb0604f3db8195cc826

Observation e29e318d-d351-4b0b-b673-5f2dafab736c · outbound

This paper cites L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.619472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.619472Z digest=sha256:2c7b78fdbb3cea72935681d60c30aec11c01d712ab0310fb3ed58687bb882b96

Observation 6c068f11-9a6e-4a59-b3d1-ebd2798c923b · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.751404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.751404Z digest=sha256:030ee5fa6ba5fec20c6713e36eba91e07d22f6677a9904388b427a10bc580e89

Observation 895298cc-c608-419f-998e-e946641413b7 · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.860327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.860327Z digest=sha256:cb7d55ba53a508521a26bf24113cd94ff0d907860b9c6f53a22def90ca40f13a

Observation a1636c54-3f43-496b-9d53-5b339f20af2f · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.997075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.997075Z digest=sha256:9d7f1c04bd2819f7ca29d1eab1e7b9573c7ec2720b4c755d38d73b6e9399a0aa

Observation bf5438fe-1fe1-4d58-925c-5c8a7a0fdcdd · outbound

This paper cites UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:58:43.660736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:33.119159Z digest=sha256:2a64452802c927a3bb24350960ef508eb8fd2c2f450e19c8977b68ddba255ecd

Observation be36cb6e-4e8f-4c7f-a43f-6d65697acb32 · outbound

This paper cites Wav2vec 2.0: Learning the structure of speech from raw audio.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Wav2vec 2.0: Learning the structure of speech from raw audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.224038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.224038Z digest=sha256:c16f36e25e8d22503f62b53ef2939c8cfd5feea216b4f2971731ea87e052b24a

Observation 27ea4246-ec6f-4fc2-9b03-b0f163153d80 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiolm: A language modeling approach to audio generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.386322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.386322Z digest=sha256:ad19436938e571b8197fa678638758f0f511b11fdad3c43c30b0318048d9d45b

Observation 7bb04f5b-c50f-4ee3-982c-7d70c3d03ece · outbound

This paper cites D., J \' u nior, A.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction D., J \' u nior, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:56.927864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:33.522304Z digest=sha256:c4d887aad1e25d3973cdf91038d4c41414486e021fc15c195392b54bead08aaa

Observation 9e0bea6d-ba96-4772-ad3b-b985ffe392f6 · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:56.691490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:33.640729Z digest=sha256:3354b5598c336121c88916b708b5e0648c01e4bb8ef4c08ce37a6264a4413647

Observation a3105a6d-23ac-44bc-adf1-d7dbaa5fd501 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.743754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.743754Z digest=sha256:d19d5fa96df1b7d6315b78cf5606cf7984734906cb319e002c0a0b15827fa66c

Observation 78a49792-6902-455d-b0a8-d4bde5038fe8 · outbound

This paper cites Qwen2-Audio Technical Report.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Qwen2-Audio Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.890548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.890548Z digest=sha256:8c26488f534989792b6c7daa0961ee309b51400267f2b25f11208cae816fb803

Observation a5a904c0-bc0b-4e51-b1b9-2c30286229e6 · outbound

This paper cites Fisher english training speech part 1 transcripts.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Fisher english training speech part 1 transcripts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:56.470785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.008369Z digest=sha256:dd0c54dd0f32968bf75e70a8faf3e8c92bd5ea9d9e0a90b005a13613dbe49c67

Observation 22a12f79-7fe4-49e1-8e7f-61dfd7ec94ef · outbound

This paper cites Simple and controllable music generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Simple and controllable music generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:56.194214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.130601Z digest=sha256:0a3bef93a4933fbecef5664e526c51576eb20108dbd24ef83eeb3a1cf96c1e85

Observation 312dfe16-cff2-4271-8325-6e2db28426ef · outbound

This paper cites Y., Ermon, S., Rudra, A., and R \' e , C.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Y., Ermon, S., Rudra, A., and R \' e , C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.926907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.241000Z digest=sha256:ecc476e470212985ad4e735e68eb62584c21545b8693dd7a5a1f765936274df6

Observation 7d474f38-9b34-4734-a226-eb1c4ca373f0 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Moshi: a speech-text foundation model for real-time dialogue

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.619710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.396487Z digest=sha256:412966e7036d7f34e17d1b2ce9be219f090dde9feee55d82670163c18248fefd

Observation 8d72e998-71f2-475c-9171-1452dc0e2651 · outbound

This paper cites Pengi: An audio language model for audio tasks.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Pengi: An audio language model for audio tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.424225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.505924Z digest=sha256:3b096f443ab61e60bcdc9670f30fc5b861b00c317a55e0aa86f9165959b9e56c

Observation 8035faf3-7c08-4760-b83f-b6f966129199 · outbound

This paper cites The Llama 3 Herd of Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:34.644720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:34.644720Z digest=sha256:f0fc357923a134e615ec89ed6ac15b1b836fde79c0cd73cc3ebe7fa65c09f40d

Observation bc2143a5-894e-4cd2-8ea6-84f5f8324f3d · outbound

This paper cites A., and Wang, H.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., and Wang, H

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:55.196143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.781530Z digest=sha256:6f25f61a6e8477e7946d21de0810e759f5f8d89792c0b6a22e20b452bc57d346

Observation 92319e28-3802-475c-9737-cb73be4d4fd7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Taming transformers for high-resolution image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.968188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:34.905358Z digest=sha256:d046b5ec5e1ebf1e61fe1a45f2bb8b218b94093937ac4e1511a986ec1ef63a29

Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:35.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:35.040733Z digest=sha256:fd249f86170e9d8cf09920b4a6aaaf065e9990922737d95ca531586922dc26c8

Observation cbe33887-8a6b-4137-8a82-948411c9230e · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiochatllama: Towards general-purpose speech abilities for llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.677786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.199376Z digest=sha256:78cca216e6b8b16f3412cbbacf40a902d2a669433f3f94f7e81466b01007f4c7

Observation 3827e764-6204-4997-b37d-b787117fdf61 · outbound

This paper cites Vita: Towards open-source interactive omni multimodal llm, 2024.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Vita: Towards open-source interactive omni multimodal llm, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.488112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.296527Z digest=sha256:90cb6bb0274cfab392f805a79c9a662201bc4c031a0207e36ad38d71ada33587

Observation 3af4459a-7331-4898-9375-ffa2d8b34895 · outbound

This paper cites Funasr: A fundamental end-to-end speech recognition toolkit.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Funasr: A fundamental end-to-end speech recognition toolkit

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:54.283084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.436272Z digest=sha256:083074e721c13cfccd681c4e9564733d52cef1e44f1b740e1abba5db67155373

Observation ac321634-5ca9-484a-8c90-13f474f18b1f · outbound

This paper cites A., Gat, I., Conneau, A., Kreuk, F., Copet, J., D \' e fossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Gat, I., Conneau, A., Kreuk, F., Copet, J., D \' e fossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.961536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.569886Z digest=sha256:3eee7411ca9c66d83000e0a62df1fd72b9f18b5cb3dad160620dcf388d9c737e

Observation 551d551f-dee4-47fd-8ee6-60bd884b59d8 · outbound

This paper cites Kenlm: Faster and smaller language model queries.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Kenlm: Faster and smaller language model queries

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.726981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.700587Z digest=sha256:b436151c039a79bf32de8e4602e8da3c12e33ff7d97fe1411e1bb2b4e5634db4

Observation 8dc67ce7-02a3-4c47-882b-a35b76ec12f8 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.570099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:35.899089Z digest=sha256:9282f2eb088d8fda5c2d35f65ee2477c91ccb18c0d6c074accaff33e013b08bf

Observation 70df1984-f0c7-4d8c-8fcf-ba11ed3134da · outbound

This paper cites Mistral 7B.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:36.013224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:36.013224Z digest=sha256:597930152e5c2455cb14c618f6ee06eab475afdf2d134bf59d93dbfd1d3408e1

Observation 3edb39cc-aa46-4f93-b2bc-0f6f5ec4b43f · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:36.096538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:36.096538Z digest=sha256:3cbdaf709c97f0791cfc55eac31e761b746826d36ec6a8d99aa4b5aa5b4792da

Observation bc7481c0-d0d8-401b-a7cb-ccaa79b20de2 · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speak, read and prompt: High-fidelity text-to-speech with minimal supervision

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.286789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.223271Z digest=sha256:ede7d87fcbb644fa4f977314829b7aaa63c1e46b3fb19145245bf4715b6c7cfe

Observation e6a44f7e-c36a-485a-8fcc-6b0313bede3d · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:53.057115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.333973Z digest=sha256:be6ca787445fe952d54973150d1fd695e196ed234ff83b0fa930db4ff8772ba7

Observation 4acb6830-e1a4-4ae2-9dd2-108b2a38de82 · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Diffwave: A versatile diffusion model for audio synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.777998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.431685Z digest=sha256:3eabd4230b81a2b4d48b64d1bb6cf3027e9c20bb8eac4e86ceec75d9a9e7c382

Observation 8aef1af1-6c15-46ed-aec2-25318c405e47 · outbound

This paper cites Audiogen: Textually guided audio generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiogen: Textually guided audio generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.568527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.542580Z digest=sha256:7c4e946efa61f73131b851e28e1f948be15a2c06aa431c116b838657d1d5e9fb

Observation 99c49b36-5fe5-49e4-9401-1637cbae224d · outbound

This paper cites Voicebox: Text-guided multilingual universal speech generation at scale.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Voicebox: Text-guided multilingual universal speech generation at scale

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.376290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.660333Z digest=sha256:ffd1e898f85187f1a1d29b52c8287ac7e61d32c0f4a8fbf936ba8013df59c988

Observation c18203e5-cdca-4886-98ff-5af9c24c8ace · outbound

This paper cites Autoregressive image generation using residual quantization.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Autoregressive image generation using residual quantization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:52.151875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.826730Z digest=sha256:d1ce54d85351f1cc4af0380c00b95bbe96218650a79c899df2e49274415abc18

Observation 9cada8b0-4ef0-4bcf-917b-c8d5e96ebc34 · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:51.953854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:36.965426Z digest=sha256:f4be8aead83d6ac5a7c54b143b21d13f7dc20252a570aa381ceb36ae5ec98282

Observation 51370c96-703f-4c0e-a66f-2859258b346f · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:51.714687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.080037Z digest=sha256:23a247ef93aa22a63909163e290b1efbc21e77038c30753f87e513b1bdb2b09a

Observation 7cd156e9-0f03-46af-80cd-2e2aa83ec3e0 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Autoregressive Image Generation without Vector Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.209661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.209661Z digest=sha256:21bfd77c72e8195e8ee4e4efb2e6164b0c803c9b5a9bafb848db505d57d25726

Observation 5b84a7b0-9973-4d37-b53d-67fb7fac0404 · outbound

This paper cites Evolutionary-scale prediction of atomic level protein structure with a language model.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Evolutionary-scale prediction of atomic level protein structure with a language model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.351317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.351317Z digest=sha256:263ad07750fb1b1dbe49465a6fe6e0e925615b0148e7cb346cc1ca442f328857

Observation 577737c9-fc2d-49ef-a029-31c956a9fa82 · outbound

This paper cites P., Wang, W., and Plumbley, M.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction P., Wang, W., and Plumbley, M

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:51.422593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.486134Z digest=sha256:45e4554a8f95d23b6bf639630d853bc75026f18789555005467acb01ab7b128e

Observation fe9c039f-14ea-439d-a24e-e7beb456a9aa · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:51.155503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.574388Z digest=sha256:c9bc3ab552aa610c197a95840eebd0fb081d1ff3f33f3b2b1dcf225980b99deb

Observation d6e50ae6-0bd2-4ffe-89d5-8c9e59471355 · outbound

This paper cites Language Model Can Listen While Speaking.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Language Model Can Listen While Speaking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.678428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.678428Z digest=sha256:8fb038d14f7b40a6e3c9e8934ce8986bacbe41aee64b16a995458bc2ab0cd920

Observation cb002973-bc68-4d7f-9ed1-0634d35b79a6 · outbound

This paper cites R., Subramanian, S., Mohr, B.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction R., Subramanian, S., Mohr, B

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.887528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.804722Z digest=sha256:bbf7b01d780e3fd6bc1983011c496db203f437ff185a1168788e4b3eeddf860e

Observation 75dea8af-703c-4fbc-ad56-db0becf0e6c3 · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.901004Z digest=sha256:1d38f36d5493b2e478c782bdd5c02726d061495bf5f9edf422efc8aabd18fc7a

Observation d6f5d090-d7d9-4f53-aac1-56f46012cc2c · outbound

This paper cites J., and Ramanovich, M.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction J., and Ramanovich, M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.643978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:37.970581Z digest=sha256:7ddb5c38e041de955dbb82ee69b6dfa48a76216443fb763b500b70ed377f523c

Observation b78aef10-d1ba-489a-aa7c-b115592f2056 · outbound

This paper cites A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., and Dupoux, E.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., and Dupoux, E

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:50.461537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.153472Z digest=sha256:cc01db15e1f7422a0d26d2a6fdfc6c2e52f9ddf8eefc795edf490446b5f28346

Observation cc02d0d4-486c-4a1a-8ae7-bc2233c0cdcd · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spirit LM: Interleaved Spoken and Written Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:38.270437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:38.270437Z digest=sha256:d1e810a2bcfc9475e131fd2a22a0c98254214373cdbc1cfbe7844827c0490ad6

Observation a8d38a4a-ce27-44c7-8e0b-09a04e20b094 · outbound

This paper cites GPT-4 Technical Report.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:38.403257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:38.403257Z digest=sha256:47d7a0aee09cc6f4c01ae3c05d93f5d363905dbaf66ac2f6532399208e7f81d3

Observation df8a2e81-3e6c-49fc-9f93-86634a4510a2 · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:50.225321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.537332Z digest=sha256:df36218bed24f7d02b928e831b054de72de857b2a395eadb6b25675688b1f91c

Observation 27299d89-e9bc-4255-8f4d-c4c55c3362ed · outbound

This paper cites S., Constant, N., Raffel, C., and Callison - Burch, C.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction S., Constant, N., Raffel, C., and Callison - Burch, C

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.986362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.665681Z digest=sha256:b5da46ad0f3013a08eaae7a809fb844b34af1899c00b7ad2a3d478d59f32b9ed

Observation f1cd5d8c-5ea6-4c03-856d-6bec1acbf740 · outbound

This paper cites Efficiently scaling transformer inference.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Efficiently scaling transformer inference

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.691407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.778622Z digest=sha256:749078dd82b334b571e4d10653cf9309cf94a2c12844198d3d00443166e7f612

Observation 42664b0f-f8b7-4927-9566-7a989af13eb5 · outbound

This paper cites Efficiently scaling transformer inference.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Efficiently scaling transformer inference

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.426335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:38.895582Z digest=sha256:37d62a7ed0fe023933a9c6dafdb21c1b8bcdfecb81a54aca3bedacd87d675efb

Observation b1d2de32-f4f6-489b-80c9-03afcfa5fba9 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:49.176260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.038734Z digest=sha256:a3bc0543be6b48dff3f1619556904402892ce5e65724f49e73b308c4dabe63d1

Observation 5bca6c3d-f9c0-4063-a333-af37c71766f6 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.949112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.145468Z digest=sha256:46ffa4a466529be97a4ffbcdf618618918d264de2dd8d1b74bcc4b610dc31430

Observation c0a44f52-f6db-4094-9d7c-9c56bafaca96 · outbound

This paper cites The candor corpus: Insights from a large multimodal dataset of naturalistic conversation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction The candor corpus: Insights from a large multimodal dataset of naturalistic conversation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.665985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.261569Z digest=sha256:90879375d59f045408b5aaf8ef4e64e28979d7021232ffa93d6d8f1a1ed1854e

Observation 7f46aba9-72bc-4c48-9fe1-a15b3e8476f3 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:39.378115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:39.378115Z digest=sha256:f541fb91d7db046ffaf6dfa2584dec99ba4883de7529d0db0cc9a917ba99d236

Observation 251546c9-78fa-4bdc-8f80-5870134b192c · outbound

This paper cites A., Bekas, C., and Lee, A.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Bekas, C., and Lee, A

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.434728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.514408Z digest=sha256:0342f13d6b04d9bbcff1f237c379fa1a80fc240d4a685e195c42cdf5cba21ad0

Observation 4be0a642-827c-40a8-89a8-0ade6f0c6960 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:48.208784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.602846Z digest=sha256:af0d8e2dd9baf685addfab5d215b2a9317373fbccc93e48553ef333ea6ec3fb6

Observation 4458dc50-c08b-4f9c-9dd6-67226ad76cde · outbound

This paper cites Graphaf: a flow-based autoregressive model for molecular graph generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Graphaf: a flow-based autoregressive model for molecular graph generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.985599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.699015Z digest=sha256:4278fdd9c68158a963715930ff3c4908830d4cfd925e853291177125794b36de

Observation 76212d70-838c-4c6e-88c9-e8652cd920d8 · outbound

This paper cites Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.748924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:39.858181Z digest=sha256:90bee0009175b44fb1461e776e3796c420889832c95e2fd066fcb44d715de0b1

Observation 0a36bdfc-2d8a-4785-9013-a3753838569f · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:39.979331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:39.979331Z digest=sha256:8e6d8af51a38a6c37feae9da1af77df049f88b1ed3dc7fff0f6b54317ec4b2bb

Observation 2f64fc8b-29c9-40b4-b397-55cc1f6fb24d · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction PandaGPT: One Model To Instruction-Follow Them All

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.100604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.100604Z digest=sha256:27a7d188cb1a072b58569a9c4e4385d1771644f451b1d1b4253a51f3ff112f96

Observation c0eecd77-e1a7-4f59-aac7-d6522552d8cb · outbound

This paper cites an unresolved cited work.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:47.513859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.217367Z digest=sha256:9d72230e6ff9f8348d57f0036e83f23f3d5be25f3e10de0782b6764362e473c0

Observation 53eefead-cdb1-4489-aac4-d01aed5c75f5 · outbound

This paper cites SALMONN: towards generic hearing abilities for large language models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction SALMONN: towards generic hearing abilities for large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.257441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.305666Z digest=sha256:422fb68ccbbd30160c4f7bede438085f78b740c1e683e1a3f03a61da7b402483

Observation ef1f5c20-0cd7-4bb2-a8a8-2266bcff8bb3 · outbound

This paper cites C., Ture, F., and Lin, J.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction C., Ture, F., and Lin, J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:47.008096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.406006Z digest=sha256:63e039202e5e18c5b5f780d31f229321cc6cec0738e5862e6c9152cf98d649e9

Observation ead31476-df3c-467a-b001-a59e2152fb08 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.519186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.519186Z digest=sha256:9a38caa7f11449be7c11f2dcbf66d01e3f74c4b0c6eda2f31dd15eb510c966b2

Observation a3628434-6402-4c47-b395-c74b301bb76f · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Gemma: Open Models Based on Gemini Research and Technology

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.614181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.614181Z digest=sha256:9e301bce750a5c83f4e7855ba0fec8ccf6bfc8700244c02d023b9f4f60eb838d

Observation 9b29c5b4-6da7-4c8d-952b-b502da87403a · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.740546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.740546Z digest=sha256:673a71a6678ee487c13ef8a4dca918e0f915249e9a65ee1c42825f76a110b3ab

Observation c6014061-9c6e-4f7c-8312-53d06888a81b · outbound

This paper cites Neural discrete representation learning.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Neural discrete representation learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.790852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:40.872224Z digest=sha256:1b0859a2a3ea9177bd785a59dedf4ca34bdcb4a7006a5c0859d5d101cd000a20

Observation 66c2f283-ffc8-4f9b-9b40-b25d6adb6bc8 · outbound

This paper cites Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.965073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:40.965073Z digest=sha256:19339762caaa482a3c6e41529e3079e21843abee1f1e32b230431378219a2dcb

Observation 4870119f-084f-4438-ab2b-8d42b446748a · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.093509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.093509Z digest=sha256:df1072fd5e9ce3150a34e012373ac9947b7ccf710c1b4f4d9c717b8609846351

Observation 30b2b3e3-7e84-42ff-a520-9b7525c41262 · outbound

This paper cites a ckstr \.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction a ckstr \

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.543674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.201189Z digest=sha256:561414b446f8e127fdf5390fe690abc81be111171266a6d87dbb38db1f14ec57

Observation feb87d25-b143-43c8-b5d0-e1124abada8c · outbound

This paper cites SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.346118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.346118Z digest=sha256:6a70448287cd227e280d0d9dcaa800a13874165f947721bbfa53e041e24e95b6

Observation 7218aea1-0087-49b5-831f-f71b3e529ccc · outbound

This paper cites Next-gpt: Any-to-any multimodal LLM.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Next-gpt: Any-to-any multimodal LLM

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.249196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.460790Z digest=sha256:9c4715ee8724530fe48099eb4c4703ded19ab489ff6b25d895a5e4518724bb4e

Observation 682e39be-d364-4755-8a56-0b5f53469a1f · outbound

This paper cites J., Wang, W., Lin, K.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction J., Wang, W., Lin, K

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.929927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.609319Z digest=sha256:927326788e9acd33dcb08b54fc05a9944848e86dd48cae3db06403407db50789

Observation b034d105-1558-44c9-8ac2-b0fa4e255f43 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.742682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.742682Z digest=sha256:bb60ab3260f7287e267dfcf1624f02fe2e765ba75defe4bb19e34646b1f7666a

Observation 5c7c0083-2de3-4e05-9bd1-c91a2c79b57a · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Diffsound: Discrete diffusion model for text-to-sound generation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.640110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:41.844624Z digest=sha256:325411e24c52eb0a765786c3cfdfe0c045ac8f694d8f379e631f147d470a61fc

Observation e3e54503-2c5e-4ae0-946b-8f85d0a409e8 · outbound

This paper cites Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.384003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.002641Z digest=sha256:1f34eac4244f79a9ae8e5917335b3a72de5da8a40cd7ceb3c95b2420821ff805

Observation 3596fa4e-db4f-4031-b596-9e8ae20c528a · outbound

This paper cites L., and Leskovec, J.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction L., and Leskovec, J

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.100642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.157901Z digest=sha256:af04707004fe6ee860c8d77c88b8c9c2324c1a0d09dde6cfe77a80e0d8380502

Observation 4e621224-e9df-44e0-abb1-1c2d5ea4649a · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Soundstream: An end-to-end neural audio codec

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.854884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.301302Z digest=sha256:96a7dd922606a6f5d7f6c0eac7daaf54ff20f33e91ed483e35d551822ef9a3dd

Observation 0ce57509-82bb-4136-938e-900155ee3eae · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.569687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.412141Z digest=sha256:ac3e67671f6b4c7a6f8f274d30fb1fec4b7463d1e15e80704418362ee030de9c

Observation f1ef71db-56d8-4cd7-bc49-777c03c8b47a · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.528594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:42.528594Z digest=sha256:956df1cd2e460b4c38fc01d87acfbc5cf017557100a6e4d77176f536df9dfecb

Observation 4a439b41-1ec4-43d5-8231-6bb96448f359 · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:44.289109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.650407Z digest=sha256:627d3f4865d391f20e757bf42a32ba23afeeeba9ad4b9a8b6ad3dfbfe9c4ff99

Observation 997ba1d2-2a75-48b7-84e9-318ae13e46b4 · outbound

This paper cites Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:43.964694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:42.812383Z digest=sha256:6ecde261401a19d22c8c29dc12067ed757985b1ba9facb2397ac9f77e22d5146

Observation 69a63b19-1c08-4ca1-928d-38e2304b0bab · outbound

This paper cites write newline.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.951304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:42.951304Z digest=sha256:04afce6159cb09b57b44ce3c07b122993717130fd0d32dca66a8ff31420f341a

Pith citing papers

Observation 0f6d85ff-3033-4893-b0eb-51c374cca2bf · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.310003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T21:22:38.888528Z digest=sha256:d259b1846559c0748f9b4d66d26e36484a9dce1f52bf8d5a44bb4b3ae6d7e11b

Observation 05a346ee-4bdb-427e-b5f7-2ae849d7bb49 · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T08:53:13.408944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:53:13.408944Z digest=sha256:08c5c428c2be3d5bd05c7caa1088423a9d8e797b52d69cb3ae8de1849bda9dcc