Pith. sign in

Paper Citation Record · LEDGER

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.11344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11344 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:57.190957Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19df71af-a646-453d-bd7f-104a648e1906 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-07T04:15:58.149537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.728975Z digest=sha256:6c7c0aa2651b299b1372239369a18a4b203e971993f65a34cec43a5bf48f49a1

Observation 993142f4-f037-4220-971b-ab252f940ef2 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:16:00.034386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:54.834097Z digest=sha256:d5ad688e6e5f3c99cce786e28e7e2c075c427969f2a14f68865f6665e343e28e

Observation 5ed8ae07-b441-4f7b-a5b7-57f8936c76ed · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:54.918846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:54.918846Z digest=sha256:f3ff427ab1dd3dd60ced79a3a69b91a7d15ed7ad25d09677448db8b01f74abd7

Observation cb01b61f-fd41-4c8f-b598-28dd8044266e · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.993991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.020157Z digest=sha256:5eaf9bc20d74db16bbd9ef98051164a6df77ba16fe0581d884ed81aa1fa1b777

Observation c994171b-17d3-42ab-91b2-5f4411e4ea5f · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.109864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.109864Z digest=sha256:2723d6e1527db16d9b98403520159fdc73dac5808db2beab88745ddffa91c5af

Observation 17be4d74-41e8-4f01-844a-18642bdeb42d · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.861434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.177313Z digest=sha256:fd179893040b7084373ff81072ca08e95471a5f06a860a3fb2475de2ab11a95d

Observation a1ed1b1b-f3e6-4de4-bb21-50ba6b6d63de · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.800209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.283442Z digest=sha256:65031185fdb4b4f218e4c468b46b4cc4c750cada3c2d14fb37c3843748eb8555

Observation ace6baef-267f-4a29-ace1-3d71ad5e36f5 · outbound

This paper cites Chafe, Charles Meyer, and Sandra A.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Chafe, Charles Meyer, and Sandra A

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.690637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.365852Z digest=sha256:f5c62936fcd5fad305c0df0c748a326f6124fb285fe01a499c60097abcb2933a

Observation 2e03906e-358a-43e6-a2ed-311dbcc09d67 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 9

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.520517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.466266Z digest=sha256:3c95a4cd02b886c9bf2c6fde9bdf0af8e843fb01f93d6989b5e1b1f0a3dc1373

Observation 1e835360-9bfd-41d2-a04c-660acea1262a · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.551726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.551726Z digest=sha256:2bb442c57afbc9a57eac3765a9aab467d3364a36cd4dc649fc4efce8ea3e4660

Observation 79edbcff-71b0-45fd-9ae9-ec794a32c25b · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.634299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.634299Z digest=sha256:03bb31c8b89251cdf0c4ddc0e9d8361f0d4abfca3eaaa886c7d089eee4422c34

Observation 56f28ec8-925f-4801-b71f-28a3d2e7124b · outbound

This paper cites Partially Observed Discrete-Time Risk-Sensitive Mean Field Games.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Partially Observed Discrete-Time Risk-Sensitive Mean Field Games

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.779217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.779217Z digest=sha256:76405583ab9a041f9e28c52cb72d1c3ca360069427cef0d14586bfdc1946181c

Observation 5bf6fa3a-dd23-4e8f-b5a0-d3ce690a832f · outbound

This paper cites Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.267143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:55.873688Z digest=sha256:5c6c98d3372da3ba80bb1ee80ee99473072386ea734c4dc4769bb1d78ee7af19

Observation 8a449894-0fcb-4aea-80f3-7b42929f58de · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:55.984530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:55.984530Z digest=sha256:4ddb248fbb43f9dda43d7282e1c0d926a734d065f4282e5fa37b81f8f4083c79

Observation 4de9f800-773f-4349-84b4-e4badef5a1ff · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.124519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.124519Z digest=sha256:b45e9ceae1c979c8fbd52235dbd5b1e2849ad5973f0069e67b762e508996ff85

Observation 080148ad-84de-4e4a-a935-70bdf49bc317 · outbound

This paper cites Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:59.038611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.218052Z digest=sha256:0962b075a3a72e675d83570ce1741fd55e180c65f18b59b699c2ef9af56e4ac4

Observation 3942f902-e29f-430b-82eb-ffb37e5f278b · outbound

This paper cites Multi-scale Speaker Diarization with Dynamic Scale Weighting.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Multi-scale Speaker Diarization with Dynamic Scale Weighting

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.855167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.335846Z digest=sha256:e50249a99ba562075a0c162d1905b15c3214b14ea193cb385f66c5cb8122a9b3

Observation af58f791-1946-4fc6-ae89-d2669e4cb1e4 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-07T04:15:57.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.441327Z digest=sha256:e76542bfeec58e370ee9818ef68cff944f275b39f644f57343df4f21b4eb0fd7

Observation 301c54e8-c851-4259-9ddd-bfd8ba67088e · outbound

This paper cites Pedregosa, G.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Pedregosa, G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.524243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.524243Z digest=sha256:2023ec3a98c0760e07bf90d7b5da4090f6f61f55974974d3d2652cb0332dc490

Observation f58fe505-57b4-4c36-866e-44a38872b049 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:59.558334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:56.604502Z digest=sha256:7f53100c969c6ef373044c5f410a893571418252f3443a45d1be8c0ada0eed7b

Observation 5092846e-ce14-438f-b405-42cea1176d02 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.689663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.689663Z digest=sha256:95976c155e107e5464fd6dec50ad39f8781a47aeca052d94b032d887342a8c00

Observation b064a4f8-1330-442d-a145-63034f2a7c02 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.769223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.769223Z digest=sha256:0a7600ccca2db2ae7ec82a9718e13d1b76011d5fe587d4d6f1b6e7159ae69514

Observation 80d5c565-14c0-4f84-9a26-578fc0102d5e · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.818718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.818718Z digest=sha256:0f0827d636e1af3fc6c64fbddb4a3e375f770a860a9de5e13c7d84df0f2ebdab

Observation 5c3e1fbd-692c-4af9-b0eb-b9c200824229 · outbound

This paper cites an unresolved cited work.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.910787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.910787Z digest=sha256:83e82b303997e103d0721f29956ab6b68eb837eeb99400afcf3fe51f59417357

Observation bb8a8c06-2f04-4733-acc7-4a8a7c0ad7f4 · outbound

This paper cites TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.644634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.027068Z digest=sha256:e73cac053230076527e647a947d3009e64fcc5ed53709c6f7d3bd3c245b8cbbc

Observation ad09241f-898f-46c1-984a-091958cd9c27 · outbound

This paper cites Speaker Diarization with LSTM.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Speaker Diarization with LSTM

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:58.310770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T04:15:57.099168Z digest=sha256:213c1ef42e084c4d5bc5ada48cddbbfccd16cb658b42c3dbb23794cd5eb4eae3

Observation a53b12da-abca-4030-807c-aecd224d6911 · outbound

This paper cites DiarizationLM: Speaker Diarization Post-Processing with Large Language Models.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:57.190957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:57.190957Z digest=sha256:deeffc20545a0a664d24ecc802039866203f773679d581766cd2c74a89664866

Pith citing papers

No inbound Pith citation observations are available.