Pith. sign in

Paper Citation Record · LEDGER

Calibrating LLM-Based Evaluator

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.13308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.13308 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:05:54.291216Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8c32b91-489c-43ac-b4fa-8f140d97ad27 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Calibrating LLM-Based Evaluator

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.204544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:9cc47adcc733b6611f8f791f4221ba94a22209134747dd1723e198622f56f9bb

Observation 2784833e-994f-4b3d-b42f-4c80756c79fb · inbound

Towards unearthing neglected climate innovations from scientific literature using Large Language Models cites this paper.

Towards unearthing neglected climate innovations from scientific literature using Large Language Models Calibrating LLM-Based Evaluator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:05:54.291216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:05:54.291216Z digest=sha256:9c8d8dd54f01e6949e578c0c6bcbfcaecb1193daf2e981829ef00dea3f980d0e

Observation a57f28fd-7676-4364-9c5f-b975c931315e · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:e6590fd94fc3f62fe6ce699da823eda364bd4d23097618dba9187f0f85fa6f85

Observation ff2cbbb3-7af3-4f8b-927d-6c210e994160 · inbound

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator cites this paper.

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Calibrating LLM-Based Evaluator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:15:44.970534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:15:44.970534Z digest=sha256:bcd46f0113ea2cf2bc161c64de153cb091ceacdb9ca9a5a94dc6b939e9b6f0a8

Observation 4ecfa734-88d1-454a-9f7f-1d66f2b609a9 · inbound

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions cites this paper.

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions Calibrating LLM-Based Evaluator

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-11T20:37:54.919951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:37:54.919951Z digest=sha256:ddd1d463c74443afdd6bddf10c74c67f145683024e08fa9c065c7f33139e862d

Observation d7b9ddb2-75e4-42d3-903a-7ffacf660847 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Calibrating LLM-Based Evaluator

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.331051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:7bf03caabf08e04605215373f1db775dea2933ff0c0327cc8feffcf14f4edc96

Observation 90b2acb0-4145-43dc-98de-32ca74b6b5e4 · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations Calibrating LLM-Based Evaluator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.260875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.260875Z digest=sha256:7bd5f6c43c47ec5e607b37c12c04c9b4a709f70ed314e8002b0e8b99805cec10

Observation 8a8ceceb-ed1b-4e88-874f-d08af2c85220 · inbound

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents cites this paper.

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents Calibrating LLM-Based Evaluator

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.894987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T09:36:50.323723Z digest=sha256:e1f1c86686c60abf9c55ec743f4ce39b1cde7adc9f1e4f262522b98a748d7c8a

Observation 5d5550e5-0e70-4a49-92f3-1286aa587f9c · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Calibrating LLM-Based Evaluator

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:55.567886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:55.567886Z digest=sha256:7e6b278cdc48e8c4d8ce618e23136f4a0c1c415cc86e9a86f39193ef04dd1a44

Observation 3e8a5565-b587-45dd-90df-fa8a37b56243 · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Calibrating LLM-Based Evaluator

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.954717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:49473d10021508e72df0a4caccfb4b89406c02d2b060855332818d75abc1a902

Observation 3ae689bf-abfe-4479-937f-c253e9ab8569 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Calibrating LLM-Based Evaluator

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.496695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:1f672f0106c3fb80c670884aad64d397aa28eef254f5ff6b48aee630a317402b