Pith. sign in

Paper Citation Record · LEDGER

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2409.06656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06656 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:28.219916Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df57b59b-7deb-4786-a949-d47497047a53 · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.219916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.219916Z digest=sha256:d4d2a952eda8b6ab23e0d54f1dea9d973478947f5f943861732a808d3570e714

Observation d29b06c6-a457-4291-9d69-02ce65b9c23a · inbound

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR cites this paper.

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:47.575403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:47.575403Z digest=sha256:8a72a765fa116fa4be63fd033cac7cf001fa85cdacce349c7c12495ed73c4c83

Observation ba6a276c-475e-4e45-ac1e-57f0e35844b8 · inbound

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations cites this paper.

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:22.701392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:42:22.701392Z digest=sha256:58a1174760aba3c26c1375e625b77cc017e0c2ee475868cd6c27e830eea14e76

Observation f059a1fd-6090-459c-b269-2647fa982092 · inbound

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering cites this paper.

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:51.725915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:51.725915Z digest=sha256:a9e91c1a786bf4843a92784fedf13b7988f542848af5270e7ba8d2d9e31237aa

Observation eab73eb0-4f5e-4f6e-807b-950354e6ce6d · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.095617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:5f50863901ec3c5506e09ae8b5e65f1358fdff1204f711bfe14aac7e4bf5d558

Observation 14d97904-3a22-40ed-920a-980f82fd1886 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.049750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:f1249f211585cced1126cbaff62d178a023b0f077924cb146ee5d724285d234d

Observation 6e02d4d5-1fba-4931-8dd9-f269d5b180e6 · inbound

Audio-Mind: An Auditable Agentic Framework for Audio Understanding cites this paper.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:25c871d318a9d537606de258c188ffdc2c4496bd92c9ec6aa743547e9734c97f

Observation 3e0f178c-107a-4681-9e80-8a7186c4cc47 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 247

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.666445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:62407e4d39a1df43c7a2a6a342876cb20e6cd337e56819e37fd25827cdc2ced6

Observation 829ccd89-c919-433d-8f19-52085df9cbf3 · inbound

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR cites this paper.

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:54:10.757568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T02:07:29.315595Z digest=sha256:202b2e48b8d55a8dfdf125745d7c070f04e8a5ee4fd07d3a268c9a0e3dae274f