Pith. sign in

Paper Citation Record · LEDGER

CASPER: A Large Scale Spontaneous Speech Dataset

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2506.00267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00267 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:48.666663Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20d2571b-8058-4693-9321-3ff0d70de7f0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CASPER: A Large Scale Spontaneous Speech Dataset Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.217794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.217794Z digest=sha256:9875b77950e0079244242791a547271180122350b243d787fda1469a660473a2

Observation a10d61bd-a6b4-4a14-93bd-853225f79f54 · outbound

This paper cites The Llama 3 Herd of Models.

CASPER: A Large Scale Spontaneous Speech Dataset The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.264067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.264067Z digest=sha256:f7aeeb9f89f94a3814b19e57c4059eb78062d2d82201d08e8c66af3fddbbfa48

Observation a5aa4fb5-fcd4-47e0-b9a3-ca9464af18d7 · outbound

This paper cites Prompted llms as chatbot modules for long open-domain conversation,.

CASPER: A Large Scale Spontaneous Speech Dataset Prompted llms as chatbot modules for long open-domain conversation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.412743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:46.337326Z digest=sha256:e950eef8e7d70a76fadafcd79463cf338f3e4e19af7412405dec43a117e408be

Observation 68702bf0-4fe6-483d-8b5a-193ef597904d · outbound

This paper cites Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations,.

CASPER: A Large Scale Spontaneous Speech Dataset Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.256684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:46.400788Z digest=sha256:3aa2938aad66911c984a17b006c40f7280d7f73d534271f747dbb3e06d6075b2

Observation 01af7eaa-2dc5-47ca-9f3f-43e76e4097e7 · outbound

This paper cites When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection.

CASPER: A Large Scale Spontaneous Speech Dataset When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:49.271166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:46.473021Z digest=sha256:fc8fe4d1ad5756a09b787d0cc1385a2c945da24866ec4f4ee9591a5c65a7da65

Observation 8f8ffbc0-0994-4e5b-8f1d-117c528aabbc · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

CASPER: A Large Scale Spontaneous Speech Dataset Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.554818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.554818Z digest=sha256:708fd7505cc1406c30f631c9409cd92ccacf0baead3e89609fac39b63d7db8af

Observation 12be38bd-e5b4-4212-8fc6-b77c2f511ef0 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CASPER: A Large Scale Spontaneous Speech Dataset CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.607323Z digest=sha256:ea67aac2500c26a879be03b63a66f7d67b8f8294c28e716da75bad7b4d2b1f59

Observation 585a613a-78b9-4a1a-a340-eb7e9c2aa412 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CASPER: A Large Scale Spontaneous Speech Dataset CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.665649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.665649Z digest=sha256:d600b0c1f9becec431ebcaf330f944ea002342329f5689f2cc19627b5b89f778

Observation b25e6881-bafd-4cab-9324-aa2c04069ef6 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

CASPER: A Large Scale Spontaneous Speech Dataset Librispeech: an asr corpus based on public domain audio books,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.738062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.738062Z digest=sha256:5fc33c47b22cb9c838e8d293e6181538cd2364921dbb9bd147ddb86518bc7726

Observation b9409cea-226f-4a71-8934-78f224063019 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

CASPER: A Large Scale Spontaneous Speech Dataset GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.821940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.821940Z digest=sha256:d2a942f7c56174ac17b2c098c13d7120835734af5e78bea4d1c8cce69c1333c3

Observation def633c0-3a31-4809-b760-03aabdc36b7c · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

CASPER: A Large Scale Spontaneous Speech Dataset Switchboard: Telephone speech corpus for research and development,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.104579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:46.894202Z digest=sha256:1518e52e90cbed96474630bfb76d9995cc4b17196d12129b61da324a9ecefaa9

Observation ebbbb0b5-1a56-43d1-8e08-a6798db13e8f · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,.

CASPER: A Large Scale Spontaneous Speech Dataset Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.962247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.962247Z digest=sha256:3d64363c29d5402ddb69ad5d7ce86b548a3fe8f0b8b16d65ff7533005f4703a4

Observation ee7048ed-e1d4-4086-98f8-c0cb7c5203c9 · outbound

This paper cites Generative spoken dialogue language modeling,.

CASPER: A Large Scale Spontaneous Speech Dataset Generative spoken dialogue language modeling,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.946083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.029630Z digest=sha256:5b2fe41886518112a046e77473425c285cf18d64ee724292c03316ac0ab7a452

Observation b0be04b9-5e03-4012-bf34-8f2746fd085a · outbound

This paper cites Audi- olm: a language modeling approach to audio generation,.

CASPER: A Large Scale Spontaneous Speech Dataset Audi- olm: a language modeling approach to audio generation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.095650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.095650Z digest=sha256:8630b813076e1eee70685cbf95f55e0649998b3c74564b232c8aef3813c7109f

Observation 6a76a570-96a6-487f-bebe-f838a94c974f · outbound

This paper cites On gener- ative spoken language modeling from raw audio,.

CASPER: A Large Scale Spontaneous Speech Dataset On gener- ative spoken language modeling from raw audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.794228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.191571Z digest=sha256:3299f34880463670eac1f4beb6468844c14da9b0e0e7b6ba0b408ef0784523d9

Observation 98b52e29-d426-4844-84cf-f880bbf562d9 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CASPER: A Large Scale Spontaneous Speech Dataset Moshi: a speech-text foundation model for real-time dialogue

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.225110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.225110Z digest=sha256:c496e8adf5364adbeeaf4e79c835e533b61eef399c059346900416196f7f4f02

Observation 1c4e6eae-9825-44c4-9fa3-70505062d41b · outbound

This paper cites Callhome american english transcripts,.

CASPER: A Large Scale Spontaneous Speech Dataset Callhome american english transcripts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.639637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.243341Z digest=sha256:71f859ac5d49e15e26a285199d2da22161a3879258d133da6a20311686ad72f5

Observation 21adb551-327c-41e2-8348-237c749781a7 · outbound

This paper cites Santa barbara corpus of spoken american english,.

CASPER: A Large Scale Spontaneous Speech Dataset Santa barbara corpus of spoken american english,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.464850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.322984Z digest=sha256:b6275cdb9e279451f442fe4a08ca86b2cd9c48a35ce14e11ad65424d9494a2dd

Observation f58fb5a0-5858-49e1-8d2f-4cc00fd0c9fe · outbound

This paper cites The hcrc map task corpus,.

CASPER: A Large Scale Spontaneous Speech Dataset The hcrc map task corpus,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.263251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.375416Z digest=sha256:8e65dd7484b8adc42f62e6a640928dec854c9a20ba30b10fe7f6d898a76fbe4d

Observation eac548f5-d474-48af-8544-dd4a6551dcd3 · outbound

This paper cites The casual conversations v2 dataset,.

CASPER: A Large Scale Spontaneous Speech Dataset The casual conversations v2 dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.101839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.439037Z digest=sha256:bc2a8df93def1144a51b5b1497bdc69810088ddeb5ea6722f312bc40a2b532e8

Observation a4d735f0-0d94-4c9b-b4b3-7b003aef0fac · outbound

This paper cites Scalable spontaneous speech dataset (SSSD): Crowdsourcing data collection to promote dialogue research,.

CASPER: A Large Scale Spontaneous Speech Dataset Scalable spontaneous speech dataset (SSSD): Crowdsourcing data collection to promote dialogue research,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.961062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.530317Z digest=sha256:56a583f9c3f9cd422f8debf8312f7177b0650584611391116c8921e77c9669a9

Observation 93663259-1ebd-45fb-abfb-c9c63adfee40 · outbound

This paper cites Silero models: pre-trained enterprise-grade stt / tts models and benchmarks,.

CASPER: A Large Scale Spontaneous Speech Dataset Silero models: pre-trained enterprise-grade stt / tts models and benchmarks,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.595510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.595510Z digest=sha256:023c5882c39b0e27d9eb6472852aa26efe7c7762c28dc41a1b4574886809620c

Observation fe169053-4210-4de0-8198-e5079bc077c8 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CASPER: A Large Scale Spontaneous Speech Dataset Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.662234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.662234Z digest=sha256:74bf7014453511a08e2d2ae5a33f63d177b742c6ba78cd0f1d3c1e8499f9b9aa

Observation c6d6d965-b0d1-4d6d-a4aa-b65e52633e82 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

CASPER: A Large Scale Spontaneous Speech Dataset Whisperx: Time-accurate speech transcription of long-form audio,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.803665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.716069Z digest=sha256:da3aa1d4b424a66d49aa5af6b105002f01b6c5dc9f9289e5160f7b5d040c8db9

Observation f5e0fc51-7c8e-4be7-92e2-c25b8c34c624 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

CASPER: A Large Scale Spontaneous Speech Dataset SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.792349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.792349Z digest=sha256:fb8085808d29e626c190b38aa2c24a6f88849347db47ae600104711bad015b9b

Observation 2a408a6f-7572-4d92-9c5c-cfc9a5a95e59 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

CASPER: A Large Scale Spontaneous Speech Dataset wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.894977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.894977Z digest=sha256:d4605a6da7eca79795f3035ec1a203f4a8588542bdeaaac37862746d5c09de9e

Observation f8027b9e-d990-41c2-9b85-4cd0b71ccdc4 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,.

CASPER: A Large Scale Spontaneous Speech Dataset pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.643117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:47.969722Z digest=sha256:050069ea58ed452159f5a9be8760e39da42737c2c692f2cd5f032e119edf516e

Observation dd6ac9b8-915a-453e-a378-98b7b56aca07 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

CASPER: A Large Scale Spontaneous Speech Dataset Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.053460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.053460Z digest=sha256:c13e6541df5effc8be1baeb4014ff6b6e731ee7290dde06077a365f2e2d99b8a

Observation 93064c8b-ec03-4c01-b10b-2f179de41d11 · outbound

This paper cites End-to-End Neural Speaker Diarization with Permutation-free Objec- tives,.

CASPER: A Large Scale Spontaneous Speech Dataset End-to-End Neural Speaker Diarization with Permutation-free Objec- tives,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.495203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:48.390178Z digest=sha256:4c741ee31b6380ae27c9a0098370241fcd5b0a11ae238b497a3595d4b80e7273

Observation 4d4c8745-b6ff-4949-98cc-6fd0a86d2b12 · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

CASPER: A Large Scale Spontaneous Speech Dataset TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.497846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.497846Z digest=sha256:ec42cdfb7424a2eaa4a00ca5fe4976cfa4b923a8a01e8954b2be7de4c77d7819

Observation b51f5b80-4fd2-4e5b-aff6-46b514b89da9 · outbound

This paper cites MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection.

CASPER: A Large Scale Spontaneous Speech Dataset MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:49.008646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:11:48.666663Z digest=sha256:bbe7ce2bf06558943c7774881bf47f6fc31a39679f0cec82e4a60e1848280c5f

Observation e119618a-71ac-450a-aeb4-c03b60e10a78 · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

CASPER: A Large Scale Spontaneous Speech Dataset NeMo: a toolkit for building AI applications using Neural Modules

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.308009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.308009Z digest=sha256:32eee2b661bd5caf3bbe01e1135a398d783f48f34870ed4f849ef2c0f83cfd86

Pith citing papers

No inbound Pith citation observations are available.