Pith. sign in

Paper Citation Record · LEDGER

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

As of 15 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2506.05062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05062 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:00.881052Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33d6d86c-72c1-42ed-a735-2c4df29cb826 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:01.026276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.860050Z digest=sha256:0938f25d72515a7b276113420afa32c06fec8b704e43b69b349904b4d6e729b9

Observation 39b7c9d9-06ab-4d2e-880d-29cf93d08887 · outbound

This paper cites Most arguments in this speech support the topic.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Most arguments in this speech support the topic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.020315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.862191Z digest=sha256:6d97c43920fb9066304c0d17d8ea21c2c9ee38d22d41c7dd440639dccd820d0e

Observation a864ccc4-1335-4394-a70e-f46e328212bf · outbound

This paper cites Pipeline-set-1.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Pipeline-set-1

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.014024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.864181Z digest=sha256:6224feb25527d01421185344c1eab680b9e14034a4fa19f102b8603125f06832

Observation 2b187ff4-7b09-43ef-b547-33a2ccbcb096 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.979680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.874699Z digest=sha256:b5576f97a3b747d713e411434a2394b5cc81853ca9364f39c2e94384cf518e77

Observation 074e1a82-a236-4b20-a8dc-06cf6f5cd173 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.973410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.876897Z digest=sha256:4693988d91ead602dcdd40fc90517a451263cedc6a9541b60c6170eafe92d77b

Observation 6f4aa0ec-40a9-488b-ac06-e7b466a42749 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.967061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.879038Z digest=sha256:022016c4cc1e1606872b2ee422235891855b85fa203505bb8f5aeb79954f5edd

Observation ab511778-c2a4-42c0-b5da-9974dc366b36 · outbound

This paper cites The argument for reform is strong.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation The argument for reform is strong

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:00.959872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.881052Z digest=sha256:a80abe15162dc6a0a324573d655f1c17ea1ae7be2fffecfb7087c60079476078

Observation b86c1e15-fc16-49e2-a1b0-13b486042ea9 · outbound

This paper cites In general, results for the CoT prompt seem to be more challenging to parse.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation In general, results for the CoT prompt seem to be more challenging to parse

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:31:01.007816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.866252Z digest=sha256:0ecf3aaffca148fe71af65f8bbfa0ac6e888dbd9db635f577f8648726b5c34a0

Observation 02956fd4-41e3-4089-85f0-214d724bc912 · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:01.000706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.868783Z digest=sha256:ad90715c38a2a007d1d0d36a879b8c8ac79329eb5aa1665602116c6f316764a1

Observation e2db4569-ce24-402a-bb11-70f45758816c · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.993866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.870745Z digest=sha256:a97d44dc712d95af0495fd7a91a50308e76a46209487cf871f73c55ed1abb3ac

Observation 575f073d-bfac-4c1e-a1c3-2fcc5205040a · outbound

This paper cites an unresolved cited work.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:00.986307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.872842Z digest=sha256:2b06147095d0573f6ec2f6af63ef1fd41af4881509b31876849e1627f0a9c664

Observation a17ae3ce-5541-47ac-8fc1-431e287fc676 · outbound

This paper cites This speech is a good opening speech for supporting the topic.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation This speech is a good opening speech for supporting the topic

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.032593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.857664Z digest=sha256:d4071bebe1b822c8ecbc1fb060a021fae42bd003c4462a3e0008bbe56467d8e8

Observation fb420f3d-ed82-4e07-ae54-8b78c3c98301 · outbound

This paper cites InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7073–7086, Online.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7073–7086, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:01.038756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:31:00.852834Z digest=sha256:0eca32d44cd0e9535f6b0ada0df67b1080d7a6f1b0e5889c42a2b62d8754f336

Observation 7d791c09-863a-4a28-9d48-9f061c101aeb · outbound

This paper cites ChatGPT as a Factual Inconsistency Evaluator for Text Summarization.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation ChatGPT as a Factual Inconsistency Evaluator for Text Summarization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:00.849793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:31:00.849793Z digest=sha256:08870e210fe231fee944fe752905d3382c43ec6715296d2a32ad97f173025a10

Observation bc470726-7043-45dd-b61e-e400e80946ce · outbound

This paper cites This speech is a good opening speech for supporting the topic.

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation This speech is a good opening speech for supporting the topic

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:00.855058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:31:00.855058Z digest=sha256:935197f106b45720534a0acb93c3fee719ea79591471da0e93d62e06ca6bc8c0

Pith citing papers

No inbound Pith citation observations are available.