Pith. sign in

Paper Citation Record · LEDGER

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2412.11449.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11449 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:00:02.820379Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:00:02.477825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:29:04.296162Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 362bc031-fc62-45fd-8ce4-a426307fae6b · outbound

This paper cites an unresolved cited work.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:00:04.076314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.467848Z digest=sha256:a29de61af248acb5e7616dfdbc77889b43cc31d8f6d0f0352dbc3c9a53731b0e

Observation bc0db306-bc06-4066-b8ba-52afd7addff8 · outbound

This paper cites For the case of speech, we use the LibriSpeech TTS dataset [27].

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music For the case of speech, we use the LibriSpeech TTS dataset [27]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:04.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.488536Z digest=sha256:007cfa29b75b8c9cfe93ee098c67dbeb2ace3f0fb0a359a9c1d180de4add8586

Observation 2e5c4b4d-1b49-4c65-9fac-592d68c1abba · outbound

This paper cites an unresolved cited work.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:00:04.011740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.494724Z digest=sha256:37877f13b5fd98aa7b30a7fcb72d5f2b5d4a6753d1e4d508eb454e915ef0910d

Observation 7729d3d9-6cc1-4807-8946-76d1527f2000 · outbound

This paper cites We are only interested in quantifying the likelihood scores for the gener- ative architecture pre-training.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music We are only interested in quantifying the likelihood scores for the gener- ative architecture pre-training

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.932322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.514037Z digest=sha256:40955e2d5c716d19bd658c45419b6905fb615314ee1c852292f18fb38eac36a0

Observation 82ae8649-7b3c-48fc-9f31-ec5132a82eae · outbound

This paper cites The proposed architecture outperforms a purely token-based model by combining continuous audio and dis- crete token-based representations.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music The proposed architecture outperforms a purely token-based model by combining continuous audio and dis- crete token-based representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.900149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.521600Z digest=sha256:591f72edcce144cfb161275b318ed35a073f30100d4584efd4587713889cdda2

Observation e785a6ed-4fb8-4a30-a4ae-af7a8410fb0d · outbound

This paper cites We thank both Google and Stanford HAI for this initiative.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music We thank both Google and Stanford HAI for this initiative

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.876394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.529451Z digest=sha256:226be80fef7c32bed4792330fa084e68a26c2bbc9b35534f02698092a05593ee

Observation f5cda316-aa4f-4aa5-852c-433e36aeb251 · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Audiolm: a language modeling approach to audio generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.749998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.605338Z digest=sha256:c8ebdd7c8626046d964be96267d3f4684ef601c7016c9627f3d8c59545de1046

Observation f665c2dd-1a20-43f4-8462-a91236f5a276 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.613792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.613792Z digest=sha256:429625ee91b6e8e0885f7c404b0c2686ea0d97109114838d9f99e68fb149184c

Observation 7658836d-6cc0-4610-a070-d02b1bfd9d85 · outbound

This paper cites Attention is all you need,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Attention is all you need,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.845190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.543785Z digest=sha256:3db511887b78c183db18c6d5ba6cb1af2fa778eca152cf1a29a1f4f32a212af1

Observation 63d715bd-9b42-419d-b9a0-c1defdcb6030 · outbound

This paper cites Language Models are Few-Shot Learners.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.561304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.561304Z digest=sha256:25babc54dc32edc7bdc312ae43b788137ffa375dca6e7baac5a738243a7f67ce

Observation b6b9e469-c7c9-4de3-8093-7581663fd5c5 · outbound

This paper cites Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.477825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.477825Z digest=sha256:6a6c89fab9784ddc3ab7a127108e9d7daf3f3f854437f7333e0ed1343f17ec53

Observation 41a5132d-32c1-4dce-aa0c-dbcdc6375438 · outbound

This paper cites A generative model for raw audio using transformer architectures,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A generative model for raw audio using transformer architectures,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.814812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.569787Z digest=sha256:7b35c91809332b5af4c8f0b83103c33cca91f8df22e726d27d75f4c0e29d7643

Observation 1948d389-e4ef-4ffa-b2d3-bdf9fdae0a7d · outbound

This paper cites A Language Model With Million Context Length For Raw Audio.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A Language Model With Million Context Length For Raw Audio

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:00:03.384865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.579250Z digest=sha256:2ed2780eac9de82b4ced474ca910fe125df8bb2f83d00ae48f7202217a243487

Observation 852147a6-0030-4b57-934d-0f67c8d6511e · outbound

This paper cites Music transformer: Generating mu- sic with long-term structure,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Music transformer: Generating mu- sic with long-term structure,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.784848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.587047Z digest=sha256:bc2c1d76da9705d517f0d23f2fa1033145a4c1803267a4a90f3e115b53f6d9e6

Observation 9cea3149-d154-4584-a557-582ba3c8dfde · outbound

This paper cites A Framework for Generative and Contrastive Learning of Audio Representations.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A Framework for Generative and Contrastive Learning of Audio Representations

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:00:03.352827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.593871Z digest=sha256:b92a74e501eb444c88345a50788e4b8deea965858f8b46fc60cfc164beadfca0

Observation d2bfc67e-447f-46ff-a2fb-684b61a9c752 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music WaveNet: A Generative Model for Raw Audio

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.704954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.704954Z digest=sha256:476b5886709d2588f1fe8687ea526b9957d24ed1d4785a5bac4150c96bd89d34

Observation 58f88346-8bc6-4d2c-94ee-27060dc810a6 · outbound

This paper cites Audio Transformers.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Audio Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.621252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.621252Z digest=sha256:12917660e6996ad931fda6811218537b59c6d043067f42c09825af86cd4cf3ec

Observation 8384468e-a393-4ee6-aca8-843f12b7a0ba · outbound

This paper cites Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.721475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.629242Z digest=sha256:ed04d0010c188bf1e5aff42a3fcc8f6d7ec6c086f0b3192286d8140f2eef3f0c

Observation bc88afda-ab7d-4059-81e1-54b6faafef48 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.692846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.637820Z digest=sha256:a270115fd7928a5c5520152668922ecc2af14397cc77530712fa0a1ddf35b835

Observation 94d345ac-4436-4f86-95d8-d32f6e75e212 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.646427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.646427Z digest=sha256:e6ea0a585304fca51d217a0898fd07c615590972b69254d719a5832eb303893e

Observation 94c9d3b6-9c77-4c60-988c-e384a9c0773a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.667660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.667660Z digest=sha256:9b5e65962e01f72e4b5a649ab5259cc30e056a6562be476e669d32bceaf93b1c

Observation 8544ca89-6801-4570-89f0-1bfd1474ca8c · outbound

This paper cites Neural Discrete Representation Learning.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Neural Discrete Representation Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.684978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.684978Z digest=sha256:df8bd5390075c253f0f33d7f21c89775428faf9b6bf16651c46d7c051ff7ad0a

Observation 79d1a1d4-5e7f-4036-af3a-d60c1ad548f5 · outbound

This paper cites an unresolved cited work.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:00:03.964088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.502571Z digest=sha256:ffe1aedfc0cec62d8523fd66442541e40c9b565a13f420723989a387c22bb46e

Observation 8aae6c97-b2d8-4cbd-a690-ec6025554102 · outbound

This paper cites Jukebox: A Generative Model for Music.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Jukebox: A Generative Model for Music

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.694978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.694978Z digest=sha256:bd60a1183e149b814e753e3b15f1b9217a3ffa07c9e9eab9ab6981f4507f9e99

Observation ef4894c0-7435-482a-88f3-22bf0f6ec3cc · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Soundstream: An end-to-end neural audio codec,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.668350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.711597Z digest=sha256:b6473cc4b51207abb9eaa15b31df8ab3f24eaa49c4d5bdc874e8a3007096236d

Observation 1a95da40-1dc8-4f35-be2e-7e394ededdab · outbound

This paper cites High Fidelity Neural Audio Compression.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music High Fidelity Neural Audio Compression

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.724873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.724873Z digest=sha256:a704b7ba2174cf44970a8125823d73d1c6d920db2110e149f9f37d1327d4fbe6

Observation d5319be5-4530-4a66-ad75-9e56b1497feb · outbound

This paper cites On generative spoken language model- ing from raw audio,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music On generative spoken language model- ing from raw audio,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.647129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.732450Z digest=sha256:64fe53c30520606f106a2738642cc91a82aef5983a2e31df4bbc1e852bcad057

Observation 86f30933-fecc-4be8-b275-a4d90c83ee36 · outbound

This paper cites textless-lib: a Library for Textless Spoken Language Processing.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music textless-lib: a Library for Textless Spoken Language Processing

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:00:03.069028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.738961Z digest=sha256:ee2b9dea689df057236298b770964e598aba3928d44e4e2153fa43b35a14b636

Observation d436e966-8d7c-4ba0-a6c8-6cc1af5cefea · outbound

This paper cites MusicLM: Generating Music From Text.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music MusicLM: Generating Music From Text

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.746067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.746067Z digest=sha256:639d1aae65ab9f8a93c2820aca466ad11ec265c1348034fef403d2eb4756bd01

Observation 29b3b38f-fa64-4225-bc32-ffba9023eec6 · outbound

This paper cites Simple and controllable music generation,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Simple and controllable music generation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.618553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.754180Z digest=sha256:fbb969ca233f1ded8f33002eb3ed52275c948bb6fc920f1d10105b4c4ddb9361

Observation e5f36c6a-aafe-4c43-b936-92cdbf36fdc6 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.762282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.762282Z digest=sha256:1df9e9f3f5b8d6940a7d44d12deb8c4819524b11bf393dd054774cb0bbb5abb6

Observation 8a96e566-92fa-4a33-85b3-7a1633467b05 · outbound

This paper cites A deep learning approach for low-latency packet loss concealment of audio sig- nals in networked music performance applications,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A deep learning approach for low-latency packet loss concealment of audio sig- nals in networked music performance applications,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.587088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.776593Z digest=sha256:e48d433ed3f009d0de582c9cabb97500cbc481a7ae4d98c32b89e4cf5ff893b1

Observation f46b34ee-6b5c-4c4e-a38b-7c4fc7a83079 · outbound

This paper cites Multi-Format Contrastive Learning of Audio Representations.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Multi-Format Contrastive Learning of Audio Representations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.786797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.786797Z digest=sha256:1e2a1f158b7fc0347c1bdc404cf4d351231290b20bd977ae4b1a04b2e7a5cc74

Observation b5095b69-910a-4388-afa3-2496334042d6 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Robust speech recognition via large-scale weak supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.793373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.793373Z digest=sha256:6cfa98d0291bbc86147340933428cf940eeef377fec50667ee333551aa6d074f

Observation aaa684bb-75b4-4b82-becf-3dc6af2a1aeb · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.799420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.799420Z digest=sha256:57583895b13705686c5bf47ca1c8e9afcb02aeaed6519a5627c24ab265317d91

Observation 86d3b56c-09a3-419a-9982-f206de098694 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Distilling the Knowledge in a Neural Network

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.805950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.805950Z digest=sha256:46b81929cbb91a9810a20b24aa2315b95dba386bc5b59b876c28c333a36f1fc2

Observation abbe882d-ecfd-4aac-b51a-9a87a5514d40 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music MiniLLM: Knowledge distillation of large language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.542992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.813971Z digest=sha256:59a27836d8299616012184f5895de1775b8200ae0262470f9fe96a2a6fedefbc

Observation 109c07dd-92fa-4297-a3bd-cefecc9a5f05 · outbound

This paper cites A simple and effective pruning approach for large lan- guage models,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A simple and effective pruning approach for large lan- guage models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.497065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T15:00:02.820379Z digest=sha256:031f68afca9cce8e8dcf859f14937ce18ba540f737812dc0a329546a1038cc62

Pith citing papers

Observation b6b9e469-c7c9-4de3-8093-7581663fd5c5 · inbound

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music cites this paper.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.477825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.477825Z digest=sha256:6a6c89fab9784ddc3ab7a127108e9d7daf3f3f854437f7333e0ed1343f17ec53

Observation 72ec0687-a5c1-4e87-a590-60a3940f4dff · inbound

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models cites this paper.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.582452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.582452Z digest=sha256:a6efd7064c53c6a4c5e95f64a012b49be45d64b59780b60f24277ede03c70ef2

Observation 0262db34-b7b9-408d-b110-e776b3ad75c3 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 130

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:04.385538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:28:58.216811Z digest=sha256:9d57aeb5704fa0c4cc3382afe5b01a305ace02353765ce31d6fa72702610bc0e