Pith. sign in

Paper Citation Record · LEDGER

CASPER: A Large Scale Spontaneous Speech Dataset

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2506.00267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00267 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:48.666663Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20d2571b-8058-4693-9321-3ff0d70de7f0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CASPER: A Large Scale Spontaneous Speech Dataset Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.217794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.217794Z digest=sha256:6b27a1c534ad3440199197a393d03bdeb40bb6ee5e31ffd1a7f12128ef8c5ee6

Observation a10d61bd-a6b4-4a14-93bd-853225f79f54 · outbound

This paper cites The Llama 3 Herd of Models.

CASPER: A Large Scale Spontaneous Speech Dataset The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.264067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.264067Z digest=sha256:deadfac203798c4166a367b9aef03314debd93f272e7b8ed218f2e333971bccd

Observation a5aa4fb5-fcd4-47e0-b9a3-ca9464af18d7 · outbound

This paper cites Prompted llms as chatbot modules for long open-domain conversation,.

CASPER: A Large Scale Spontaneous Speech Dataset Prompted llms as chatbot modules for long open-domain conversation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.412743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:46.337326Z digest=sha256:f334ca7134a8aa0ac831432dd778db4defefe98884b160abf289e441312f28bc

Observation 68702bf0-4fe6-483d-8b5a-193ef597904d · outbound

This paper cites Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations,.

CASPER: A Large Scale Spontaneous Speech Dataset Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.256684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:46.400788Z digest=sha256:be8784b073db3d968bea8918a68f31a5f5f56e7bedd3c05a26950232432c1504

Observation 01af7eaa-2dc5-47ca-9f3f-43e76e4097e7 · outbound

This paper cites When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection.

CASPER: A Large Scale Spontaneous Speech Dataset When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:49.271166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:46.473021Z digest=sha256:77a70f27da3b76d10a26b813622ec7243a201fe7e204129fab6a393e02a1f1ca

Observation 8f8ffbc0-0994-4e5b-8f1d-117c528aabbc · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

CASPER: A Large Scale Spontaneous Speech Dataset Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.554818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.554818Z digest=sha256:fdb66035a5de3d29a592833b7b6df62b4c4fd65e675097dfd2c776c2ba9bbe8a

Observation 12be38bd-e5b4-4212-8fc6-b77c2f511ef0 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CASPER: A Large Scale Spontaneous Speech Dataset CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.607323Z digest=sha256:813b9161011883b3ceee1cf68052dc282f579fe07fb2d9425e6acea9db91b5d4

Observation 585a613a-78b9-4a1a-a340-eb7e9c2aa412 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CASPER: A Large Scale Spontaneous Speech Dataset CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.665649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.665649Z digest=sha256:43b8c5455efc90967667a05e06735a5f51f1996176928278f14f55b06ef5f4a6

Observation b25e6881-bafd-4cab-9324-aa2c04069ef6 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

CASPER: A Large Scale Spontaneous Speech Dataset Librispeech: an asr corpus based on public domain audio books,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.738062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.738062Z digest=sha256:80a1770f47cdf907ccc0eb9100273213ade10ccb6f4ca19b3cd555990944a289

Observation b9409cea-226f-4a71-8934-78f224063019 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

CASPER: A Large Scale Spontaneous Speech Dataset GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.821940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.821940Z digest=sha256:8231beababc7575b6734d975813f431af4544ce453398813ac82232ad45f771d

Observation def633c0-3a31-4809-b760-03aabdc36b7c · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

CASPER: A Large Scale Spontaneous Speech Dataset Switchboard: Telephone speech corpus for research and development,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.104579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:46.894202Z digest=sha256:620d83b981e10d9d1b756e0b81ef298d7aea3aaf3c24f5b7d64ea7a5e1348997

Observation ebbbb0b5-1a56-43d1-8e08-a6798db13e8f · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,.

CASPER: A Large Scale Spontaneous Speech Dataset Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.962247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.962247Z digest=sha256:df30cf2d1522564fd6ee6284a8fb72a72992751b79a80b74758da7a7ed399bbe

Observation ee7048ed-e1d4-4086-98f8-c0cb7c5203c9 · outbound

This paper cites Generative spoken dialogue language modeling,.

CASPER: A Large Scale Spontaneous Speech Dataset Generative spoken dialogue language modeling,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.946083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.029630Z digest=sha256:67722bc16fe2eb72a90956f2cd2628fae213ecc7b2dae050e12e28a780d1f7db

Observation b0be04b9-5e03-4012-bf34-8f2746fd085a · outbound

This paper cites Audi- olm: a language modeling approach to audio generation,.

CASPER: A Large Scale Spontaneous Speech Dataset Audi- olm: a language modeling approach to audio generation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.095650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.095650Z digest=sha256:708bd2493bd5a12ec5659d91121b85ed4f2c57764a5199ced7c0c5f1b4b82e74

Observation 6a76a570-96a6-487f-bebe-f838a94c974f · outbound

This paper cites On gener- ative spoken language modeling from raw audio,.

CASPER: A Large Scale Spontaneous Speech Dataset On gener- ative spoken language modeling from raw audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.794228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.191571Z digest=sha256:d599211e980be8736eb473cad3ac719a486dc659bf8731f735d3be89df63eed2

Observation 98b52e29-d426-4844-84cf-f880bbf562d9 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CASPER: A Large Scale Spontaneous Speech Dataset Moshi: a speech-text foundation model for real-time dialogue

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.225110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.225110Z digest=sha256:f88c0e81b1f66533a17d587fcb09e767d991571ed3fc518a08c6ac3af29927fd

Observation 1c4e6eae-9825-44c4-9fa3-70505062d41b · outbound

This paper cites Callhome american english transcripts,.

CASPER: A Large Scale Spontaneous Speech Dataset Callhome american english transcripts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.639637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.243341Z digest=sha256:cac3e12328ebe2f3d67ec287608d7cacecc9564bbf9dabbd461b22a6799f2539

Observation 21adb551-327c-41e2-8348-237c749781a7 · outbound

This paper cites Santa barbara corpus of spoken american english,.

CASPER: A Large Scale Spontaneous Speech Dataset Santa barbara corpus of spoken american english,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.464850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.322984Z digest=sha256:55ceaf17f7880103b2b06567fdd0dbb8534a7fa6a15ee7ecbdd0b18fe9fe5fec

Observation f58fb5a0-5858-49e1-8d2f-4cc00fd0c9fe · outbound

This paper cites The hcrc map task corpus,.

CASPER: A Large Scale Spontaneous Speech Dataset The hcrc map task corpus,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.263251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.375416Z digest=sha256:8f8b7937ef17554be94229395582489ea2d70120b509ef119616d25f1216ccb7

Observation eac548f5-d474-48af-8544-dd4a6551dcd3 · outbound

This paper cites The casual conversations v2 dataset,.

CASPER: A Large Scale Spontaneous Speech Dataset The casual conversations v2 dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.101839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.439037Z digest=sha256:3f301a1763066f5719aa1010320459cee3189f503c4293574cff9910befa061d

Observation a4d735f0-0d94-4c9b-b4b3-7b003aef0fac · outbound

This paper cites Scalable spontaneous speech dataset (SSSD): Crowdsourcing data collection to promote dialogue research,.

CASPER: A Large Scale Spontaneous Speech Dataset Scalable spontaneous speech dataset (SSSD): Crowdsourcing data collection to promote dialogue research,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.961062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.530317Z digest=sha256:1af1931d7b616fc24b73b15478282be364b3a27ef7a68933c9fef979ddbb5b58

Observation 93663259-1ebd-45fb-abfb-c9c63adfee40 · outbound

This paper cites Silero models: pre-trained enterprise-grade stt / tts models and benchmarks,.

CASPER: A Large Scale Spontaneous Speech Dataset Silero models: pre-trained enterprise-grade stt / tts models and benchmarks,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.595510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.595510Z digest=sha256:b77a5b8c26333e3de1a378ede540c56d32857697cff1ff1b6e6ee9673389edf0

Observation fe169053-4210-4de0-8198-e5079bc077c8 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CASPER: A Large Scale Spontaneous Speech Dataset Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.662234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.662234Z digest=sha256:6ac1814561f4edf2fec0be8f42544cfd9aee0cec1ae7ae4fdb650131ae3598cd

Observation c6d6d965-b0d1-4d6d-a4aa-b65e52633e82 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

CASPER: A Large Scale Spontaneous Speech Dataset Whisperx: Time-accurate speech transcription of long-form audio,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.803665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.716069Z digest=sha256:aa419926c5ad98663cd0c2ecb578e5a5b20cc136fc3a6b555acc617c57c2c1c1

Observation f5e0fc51-7c8e-4be7-92e2-c25b8c34c624 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

CASPER: A Large Scale Spontaneous Speech Dataset SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.792349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.792349Z digest=sha256:548546c1385f4ab66ae03efa3c8d9d71625537d1c3bf1104bc6a60dd775e392e

Observation 2a408a6f-7572-4d92-9c5c-cfc9a5a95e59 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

CASPER: A Large Scale Spontaneous Speech Dataset wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.894977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.894977Z digest=sha256:000b93fa4492ede795903bb348c9199f372f81cb14e2a1955d1daa58795c8d67

Observation f8027b9e-d990-41c2-9b85-4cd0b71ccdc4 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,.

CASPER: A Large Scale Spontaneous Speech Dataset pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.643117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:47.969722Z digest=sha256:1a9dd6fb10e6d67d0a1a927eb25618183f6201c82f1f0b97c87b7b8cec74f9f5

Observation dd6ac9b8-915a-453e-a378-98b7b56aca07 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

CASPER: A Large Scale Spontaneous Speech Dataset Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.053460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.053460Z digest=sha256:a37904647e710625bf7cd843254a69baea803d951eb77b900791def51cf77e6b

Observation 93064c8b-ec03-4c01-b10b-2f179de41d11 · outbound

This paper cites End-to-End Neural Speaker Diarization with Permutation-free Objec- tives,.

CASPER: A Large Scale Spontaneous Speech Dataset End-to-End Neural Speaker Diarization with Permutation-free Objec- tives,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.495203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:48.390178Z digest=sha256:3893c4ab3041a969d28e68bcafd4e59acdbcf2cad87c8e9e2aaeef867d2381b1

Observation 4d4c8745-b6ff-4949-98cc-6fd0a86d2b12 · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

CASPER: A Large Scale Spontaneous Speech Dataset TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.497846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.497846Z digest=sha256:d207ed0ef424db6afe169b221b172308ae652e8167324720ec3a37c19ab2707b

Observation b51f5b80-4fd2-4e5b-aff6-46b514b89da9 · outbound

This paper cites MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection.

CASPER: A Large Scale Spontaneous Speech Dataset MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:49.008646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:11:48.666663Z digest=sha256:de415360a9cb07228b97344b42c83cfaa880d026d6a32bd997861ab560aeeefc

Observation e119618a-71ac-450a-aeb4-c03b60e10a78 · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

CASPER: A Large Scale Spontaneous Speech Dataset NeMo: a toolkit for building AI applications using Neural Modules

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.308009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.308009Z digest=sha256:66b0848ca73db578b62110429f07c08d5461584ee973b4726c72c7dff9af9511

Pith citing papers

No inbound Pith citation observations are available.