Pith. sign in

Paper Citation Record · LEDGER

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2409.06656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06656 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:28.219916Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df57b59b-7deb-4786-a949-d47497047a53 · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.219916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.219916Z digest=sha256:d4d2a952eda8b6ab23e0d54f1dea9d973478947f5f943861732a808d3570e714

Observation d29b06c6-a457-4291-9d69-02ce65b9c23a · inbound

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR cites this paper.

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:47.575403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:47.575403Z digest=sha256:8a72a765fa116fa4be63fd033cac7cf001fa85cdacce349c7c12495ed73c4c83

Observation ba6a276c-475e-4e45-ac1e-57f0e35844b8 · inbound

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations cites this paper.

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:22.701392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:42:22.701392Z digest=sha256:728f93101a36a607b9d2f5776bf9b5e5a1c9dd70ef73b6d3d41a09605fe016e0

Observation f059a1fd-6090-459c-b269-2647fa982092 · inbound

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering cites this paper.

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:51.725915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:51.725915Z digest=sha256:a9e91c1a786bf4843a92784fedf13b7988f542848af5270e7ba8d2d9e31237aa

Observation eab73eb0-4f5e-4f6e-807b-950354e6ce6d · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.095617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:465d1a8dbed066970f5397eb32f98a3a96a8f8fae3a06772af29deab38abf1ae

Observation 14d97904-3a22-40ed-920a-980f82fd1886 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.049750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:e5599cffbf13bb78bf851d0046f343d5e2f985cea7207fdcdb8b76a5fdf4635c

Observation 6e02d4d5-1fba-4931-8dd9-f269d5b180e6 · inbound

Audio-Mind: An Auditable Agentic Framework for Audio Understanding cites this paper.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:01d34d395f61f824ee160f9cdb3be73a6227255cf4cc25c34b444a49b10ce6d4

Observation 3e0f178c-107a-4681-9e80-8a7186c4cc47 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 247

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.666445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:654a1bf15e07d7be0e2491b1b05a642d5cced53418beb81b4ccbad34d828da0f

Observation 829ccd89-c919-433d-8f19-52085df9cbf3 · inbound

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR cites this paper.

Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:54:10.757568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T02:07:29.315595Z digest=sha256:4a02763f5a184264a64b853938253fdb91acd26cc921f7e9d18431fcf411a367