Pith. sign in

Paper Citation Record · LEDGER

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.11344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11344 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:57.190957Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19df71af-a646-453d-bd7f-104a648e1906 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-07T04:15:58.149537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.728975Z digest=sha256:3866bff01ba427ddb8dda32d7cb225b632e8214179dd37c23f68b89cfdf3f0a5

Observation 993142f4-f037-4220-971b-ab252f940ef2 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:16:00.034386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.834097Z digest=sha256:23e7e987406562f2614b277a45804dcbdcb7f83f48217f84e5dc621a3aa88358

Observation 5ed8ae07-b441-4f7b-a5b7-57f8936c76ed · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:54.918846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:54.918846Z digest=sha256:a629d63bd057b57896c560cde810178973cefe39191181412e4f74b51d148e70

Observation cb01b61f-fd41-4c8f-b598-28dd8044266e · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.993991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.020157Z digest=sha256:7b458f976f904c459a3f764b4947ad6d5c6041bc4f6bde10300a34b63c2ea8ec

Observation c994171b-17d3-42ab-91b2-5f4411e4ea5f · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.109864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.109864Z digest=sha256:2d4354776e753489666bb1796a90cc1e3631dc46a9de53ade5c3b5ab85b3d719

Observation 17be4d74-41e8-4f01-844a-18642bdeb42d · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.861434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.177313Z digest=sha256:108dbb206014bef788effd316235370bd11284ab1c0011a85a6035d5df43538d

Observation a1ed1b1b-f3e6-4de4-bb21-50ba6b6d63de · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.800209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.283442Z digest=sha256:12b48f2c02e4efb1289a200718472254efe49bbc2e86a376339f35cffaf7b976

Observation ace6baef-267f-4a29-ace1-3d71ad5e36f5 · outbound

This paper cites Chafe, Charles Meyer, and Sandra A.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Chafe, Charles Meyer, and Sandra A

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.690637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.365852Z digest=sha256:381f32de19165ec741c51971882c4148610dd4c32a3845b5352e6607b3aa4e0a

Observation 2e03906e-358a-43e6-a2ed-311dbcc09d67 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 9

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.520517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.466266Z digest=sha256:0b27c819340490bdc2b6f4072782d19df1029476b19324acabc1e4104ced27ad

Observation 1e835360-9bfd-41d2-a04c-660acea1262a · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.551726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.551726Z digest=sha256:0cb7ebdebaeb7807ae8e135f897e1dd91243883e5fefb68692fd50131aa668b3

Observation 79edbcff-71b0-45fd-9ae9-ec794a32c25b · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.634299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.634299Z digest=sha256:9b6bcf6c88236cfb00871f6323a43329c6a861fb117c4b40fcbf2694f7da012e

Observation 56f28ec8-925f-4801-b71f-28a3d2e7124b · outbound

This paper cites Partially Observed Discrete-Time Risk-Sensitive Mean Field Games.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Partially Observed Discrete-Time Risk-Sensitive Mean Field Games

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.779217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.779217Z digest=sha256:0e2a051a480cda5d728ac86f4a475e0ec7763f9af6b9fdfee67c067b198ac00b

Observation 5bf6fa3a-dd23-4e8f-b5a0-d3ce690a832f · outbound

This paper cites Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.267143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.873688Z digest=sha256:c457989b4841a5843f374164fa51e5e9cc7c4333a4cb1fa1a6c0324a226dd7d3

Observation 8a449894-0fcb-4aea-80f3-7b42929f58de · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.984530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.984530Z digest=sha256:2453b282018d3bdf55821d3218d96e88d71a569202bea891afbd9245617430fc

Observation 4de9f800-773f-4349-84b4-e4badef5a1ff · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.124519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.124519Z digest=sha256:bd6e26cd3501377d738768a3e7a6aba53057c178bb43acb1d66520b7bcf8ae6a

Observation 080148ad-84de-4e4a-a935-70bdf49bc317 · outbound

This paper cites Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.038611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.218052Z digest=sha256:30856003414387b11beaefdbc183af32aa017af9058f3819136812383ae4edc7

Observation 3942f902-e29f-430b-82eb-ffb37e5f278b · outbound

This paper cites Multi-scale Speaker Diarization with Dynamic Scale Weighting.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Multi-scale Speaker Diarization with Dynamic Scale Weighting

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.855167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.335846Z digest=sha256:937bde8c90853495e4dc828e80b5badc51797229dd75bad7f60ec9292ce014fc

Observation af58f791-1946-4fc6-ae89-d2669e4cb1e4 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.441327Z digest=sha256:20844986d6fb0e00f69b985a7a5b6741b2488440a106aba9917e68cd4784e6be

Observation 301c54e8-c851-4259-9ddd-bfd8ba67088e · outbound

This paper cites Pedregosa, G.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Pedregosa, G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.524243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.524243Z digest=sha256:2e79d1913603ac7053634959c06601776286847af386657050f7fccf57e2e046

Observation f58fe505-57b4-4c36-866e-44a38872b049 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.558334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.604502Z digest=sha256:ab0317e3b5f447bc3b4ac44332e63ff4713edb8461ab0e02d20dd3044d3a48d8

Observation 5092846e-ce14-438f-b405-42cea1176d02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.689663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.689663Z digest=sha256:701b0f394d2b98ef4abe6abf423a4fd8406d45c0d7f55bb7b7756f627b63d055

Observation b064a4f8-1330-442d-a145-63034f2a7c02 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.769223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.769223Z digest=sha256:39342b386596c0513eec258ecf1145cfe96c6e4b3a09c6b2a8a561fc861b9a59

Observation 80d5c565-14c0-4f84-9a26-578fc0102d5e · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.818718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.818718Z digest=sha256:6a1439e2c16c3943834c7f7c041a1ba8dcb435335a7b46d531cfb026ab32e27e

Observation 5c3e1fbd-692c-4af9-b0eb-b9c200824229 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.910787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.910787Z digest=sha256:5d5cceb6dd6aff510d5ac7f392b0a0c2e67d469bbd9bed44edbde14dfec8433e

Observation bb8a8c06-2f04-4733-acc7-4a8a7c0ad7f4 · outbound

This paper cites TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.644634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.027068Z digest=sha256:67419491b582a6bee15bc3f4e96d0ec2ee143ac2e505ede24ef62f0498d2d20b

Observation ad09241f-898f-46c1-984a-091958cd9c27 · outbound

This paper cites Speaker Diarization with LSTM.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Speaker Diarization with LSTM

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.310770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.099168Z digest=sha256:434da98fa3adbc7da641a12783c4644a160af8161c28484c17f978af20fed384

Observation a53b12da-abca-4030-807c-aecd224d6911 · outbound

This paper cites DiarizationLM: Speaker Diarization Post-Processing with Large Language Models.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:57.190957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:57.190957Z digest=sha256:19dfaa5cccf34d676d4aa482de3981297c2a8b1ed98d2f4593ed866053f09322

Pith citing papers

No inbound Pith citation observations are available.