Pith. sign in

Paper Citation Record · LEDGER

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

As of 20 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 5 inbound Pith citation observations for arXiv:2412.00265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00265 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:38:52.776117Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:08:20.924375Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:52:42.244532Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f5c8314-bbe9-4a81-a4d7-f825b5eac8c6 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:38:52.683006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:38:52.683006Z digest=sha256:5c4eec06ca81bf3df7c213576d47b44efe39e9f4822296222361e55df0361146

Observation 6a3f55e1-c979-41c4-8436-fb8231b5f687 · outbound

This paper cites For instance, in ˆH2, we observe a sequence of upper lip elevation, lower lip elevation, and finally, tongue dorsum elevation.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies For instance, in ˆH2, we observe a sequence of upper lip elevation, lower lip elevation, and finally, tongue dorsum elevation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:53.133317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.692663Z digest=sha256:6e565793b13114663f4289b5355e2da9b2b227b2bff38b0e706d4253557dea17

Observation 26ca37eb-4a74-4574-bade-863b7f6c8242 · outbound

This paper cites word": "<word>.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies word": "<word>

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:53.083317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.709596Z digest=sha256:41864e739f7b19d334c6d9879de14850d48e162efd375fd75b226340fccc0278

Observation a6f7298c-a5fc-4dba-8ac7-9be6ef0f32b5 · outbound

This paper cites word": "Hello.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies word": "Hello

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:53.017862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.731331Z digest=sha256:1e4dfe0aa63c2abb1403149f5e13ac8682495b6b6d702f6a5beb89b9fec5b898

Observation 072ec45b-5bbb-4a3c-9264-8f5d334baa86 · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:53.116084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.698526Z digest=sha256:c62142792368d034608a9c7b7ee5527b1ee6340a735e776b38c0f4c7471cae40

Observation 8f518cf7-da55-416b-a5ba-3eba14a4ae5a · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:53.100084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.704377Z digest=sha256:ab09c70cd3df948ad3583f762e52e09ea29a451678e0d1458266caad508e7bba

Observation c39445cd-4cdc-407d-a638-e839f1b0a5ac · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:53.067380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.715078Z digest=sha256:6d04cce3a8a3ce053fbf6a4d8369ea035a6a5fc65c48c4e4f7fe3ee7bf2c9da5

Observation 775c1a0e-fa99-42dd-822f-5e2d701e58bf · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:53.051530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.720249Z digest=sha256:3b3cffb04d54f334aa196f2a80767d45d55937287a94dece31b6bf2800bcefa1

Observation 6829a6cb-11f5-4df0-a4ee-c71ad6731b01 · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:53.035050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.725316Z digest=sha256:9643692e64adaabc9f363d564bbb0316f07252a959c75c26dc937cc506efb82b

Observation 5bc55373-a9e3-466c-9236-1e4dddd3375e · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:52.983114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.744533Z digest=sha256:e311e6e5356c83fb1bcf1a2124334e529f3abe936483c30092c725a643265341

Observation 4d3bb387-3e67-423c-9435-707cc257054d · outbound

This paper cites an unresolved cited work.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:38:52.964814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.750168Z digest=sha256:e968fdb68a8e6766b66e36b269bb23ea097c63c9e07b8446c2bea9f869c40588

Observation 7085cb48-d9eb-431e-9a36-9fd8f133a700 · outbound

This paper cites • Set has dysfluency to 0 if no entry is generated.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies • Set has dysfluency to 0 if no entry is generated

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:52.946024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.755353Z digest=sha256:47977fa35322ac1af1105b05b3ecf5e0d0d3e226dea763b77efc753fefe872c6

Observation 120f7a9d-3f1e-4b9a-9347-cda024e9e6cf · outbound

This paper cites word": "I.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies word": "I

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:52.927836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.760359Z digest=sha256:8519c1b37bbcc0acc88544b409e96962f9f359986bacd5d620db25060449bcd7

Observation 2f9c8be5-dc25-478b-ac54-70c4ded3a064 · outbound

This paper cites has_dysfluency.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies has_dysfluency

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:52.909620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.765273Z digest=sha256:bc34e6ebc58c9216b7bbf4e09910a93d27a731c9e4b3f9a0221ad9d5febc70da

Observation 9c9799ae-f462-4a8a-b08d-2bb7334b15cd · outbound

This paper cites The current mainstream methods treat this problem as a time-based object detection task (Lian et al., 2023c; Lian & Anumanchipalli, 2024; Zhou et al., 2024b;a; Lian et al., 2024).

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies The current mainstream methods treat this problem as a time-based object detection task (Lian et al., 2023c; Lian & Anumanchipalli, 2024; Zhou et al., 2024b;a; Lian et al., 2024)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:52.876599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.776117Z digest=sha256:f962dd5bb1546cc93d26e1bc61bcadfe38fda932e3c81771a70ce9e530782df2

Observation efd6abfd-cd2d-4750-be9e-c12f9cb61470 · outbound

This paper cites word": "<word>.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies word": "<word>

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:38:53.000341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.737923Z digest=sha256:6688c869f27118359aae49e096a8024518a3c6907c9062e18d40deb972615bc8

Observation 1122258b-5a95-4369-854e-f7d39ae76c51 · outbound

This paper cites Coding Speech through Vocal Tract Kinematics.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Coding Speech through Vocal Tract Kinematics

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-12T05:38:52.665262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:38:52.665262Z digest=sha256:ea345a829a2c7a94cf1577c7374f86b892d338642190f8088ca1ec3dee7cef32

Observation 86a10e22-b056-482c-86ac-bee8bc1bc36d · outbound

This paper cites Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T05:38:52.841536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.673856Z digest=sha256:a6241402cda0f6570d2351443b516d7a9551c7beb5e3d3f477a591076c14ec68

Observation 0104c436-4e44-4e6d-9887-a369f5cce865 · outbound

This paper cites VCTK includes speech data uttered by 109 native speakers of English with various accents.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies VCTK includes speech data uttered by 109 native speakers of English with various accents

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T05:38:52.892668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T05:38:52.770475Z digest=sha256:1ed8c18450922bed718a6a1b8de1bc6849a280dd7a2ac1f9a211aab1fb72ade5

Pith citing papers

Observation 6ad04381-2d1e-4e32-80e6-9805e4666fdc · inbound

Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection cites this paper.

Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:20.924375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:20.924375Z digest=sha256:97c67e583b3b79beef3ba65629db32b81fceb503ae22acd52518ab6b4c2ab92b

Observation d6e72374-bada-4d24-90a0-ed0d6b78a0a4 · inbound

Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection cites this paper.

Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:20.896344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:20.896344Z digest=sha256:d48794cc2b3a5e69d1d1518dc55eecac158e93b69e2d93d763cd89f242a676ba

Observation ee6b3803-d98a-488f-8356-8ee8afb308da · inbound

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis cites this paper.

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:08.947796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:08.947796Z digest=sha256:6305cbf31461c76d79814b06cd363d2b3e37d1d2dec25f0fdb0a1cd29610b4de

Observation 02cb38b8-b50b-4f95-954f-51845d3705e2 · inbound

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling cites this paper.

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:11:25.165013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:11:25.165013Z digest=sha256:3a55f9044327a8af0293925d090acf42494921b07d5baa87b24b6b98f30eca23

Observation efa9ab3e-fc93-468d-9ff8-51ff24db52fe · inbound

Revisiting Rule-Based Stuttering Detection: A Comprehensive Analysis of Interpretable Models for Clinical Applications cites this paper.

Revisiting Rule-Based Stuttering Detection: A Comprehensive Analysis of Interpretable Models for Clinical Applications SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:52:42.332945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:52:38.516956Z digest=sha256:5a3b942a0dd9ea6fdbb4ca4e38022c6a211925cdc5f28b6df605afed7bf64d0b