Pith. sign in

Paper Citation Record · LEDGER

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2506.05062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05062 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:00.881052Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33d6d86c-72c1-42ed-a735-2c4df29cb826 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:01.026276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.860050Z digest=sha256:d55093247a31249116e49183be18289178a92b61742016b853d269485b46d59e

Observation 39b7c9d9-06ab-4d2e-880d-29cf93d08887 · outbound

This paper cites Most arguments in this speech support the topic.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Most arguments in this speech support the topic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.020315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.862191Z digest=sha256:149baa9402bbb6f91e8cac7ed8fcf5090f728c37e8e71292b7542d306bb18241

Observation a864ccc4-1335-4394-a70e-f46e328212bf · outbound

This paper cites Pipeline-set-1.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Pipeline-set-1

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.014024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.864181Z digest=sha256:49dd47732d3705751c71e64726d769e9a0f1c9913701bdba63d6b38956ad78e1

Observation 2b187ff4-7b09-43ef-b547-33a2ccbcb096 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.979680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.874699Z digest=sha256:5e4c30ea8423af9ada00f233c81ed1b668b54dc28fc991a0dd97abd2bd33fd06

Observation 074e1a82-a236-4b20-a8dc-06cf6f5cd173 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.973410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.876897Z digest=sha256:f87eb478cff24ba19d66edd5882f8a552ffab608fd072de209285e79e0395d67

Observation 6f4aa0ec-40a9-488b-ac06-e7b466a42749 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.967061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.879038Z digest=sha256:92467787c058e376547e12cb013b850c27582d64b91b62444e54c6555043f2ba

Observation ab511778-c2a4-42c0-b5da-9974dc366b36 · outbound

This paper cites The argument for reform is strong.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation The argument for reform is strong

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:00.959872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.881052Z digest=sha256:e0f21ffea374fee8bfa5bc2b6c551d1ef3dad5a2e3171a8b77445c5aa0097416

Observation b86c1e15-fc16-49e2-a1b0-13b486042ea9 · outbound

This paper cites In general, results for the CoT prompt seem to be more challenging to parse.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation In general, results for the CoT prompt seem to be more challenging to parse

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:31:01.007816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.866252Z digest=sha256:13d8d902c35ae7a932a741bf0aff111ec2857c34310abe3e4d8a8c8ab89d084a

Observation 02956fd4-41e3-4089-85f0-214d724bc912 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:01.000706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.868783Z digest=sha256:800da1177508d91e26825e05ed24164b3fee48fb043d69d5830cccb4c81185a0

Observation e2db4569-ce24-402a-bb11-70f45758816c · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.993866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.870745Z digest=sha256:9cb7e1ce7df5b05b4a3710c4116b9621e8c064bcdc9c9a870144c6248d210aab

Observation 575f073d-bfac-4c1e-a1c3-2fcc5205040a · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.986307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.872842Z digest=sha256:14ab341b181b4a4eef3e317a509c7cc0a4450b18d08e34e9f65c0f73028fc2fc

Observation a17ae3ce-5541-47ac-8fc1-431e287fc676 · outbound

This paper cites This speech is a good opening speech for supporting the topic.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation This speech is a good opening speech for supporting the topic

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.032593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.857664Z digest=sha256:d0e8dbeac36d372d551808093c2584d7dc89045ccdaaff9c828c57a66cb973a3

Observation fb420f3d-ed82-4e07-ae54-8b78c3c98301 · outbound

This paper cites InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7073–7086, Online.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7073–7086, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.038756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:31:00.852834Z digest=sha256:77c912f4223b082a4e114da313dd23d41e122bd5a98e5a9b0682afbc557b9ab7

Observation 7d791c09-863a-4a28-9d48-9f061c101aeb · outbound

This paper cites ChatGPT as a Factual Inconsistency Evaluator for Text Summarization.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation ChatGPT as a Factual Inconsistency Evaluator for Text Summarization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:00.849793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:31:00.849793Z digest=sha256:14bc6f7e30d0d2c772a814ff481322f1497d0d7644745c59a9b69bfbe3f898c6

Observation bc470726-7043-45dd-b61e-e400e80946ce · outbound

This paper cites This speech is a good opening speech for supporting the topic.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation This speech is a good opening speech for supporting the topic

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:00.855058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:31:00.855058Z digest=sha256:ee2ce9f1064880511828c49ae5041724acf1f4bff05164231f7fc72fc9eb3daf

Pith citing papers

No inbound Pith citation observations are available.