Pith. sign in

Paper Citation Record · LEDGER

Calibrating LLM-Based Evaluator

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.13308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.13308 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:05:54.291216Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8c32b91-489c-43ac-b4fa-8f140d97ad27 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Calibrating LLM-Based Evaluator

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.204544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:ff091eab75fd833f3f8c6ab7e443519fff1ce3aace11eafacbedfbf306657de5

Observation 2784833e-994f-4b3d-b42f-4c80756c79fb · inbound

Towards unearthing neglected climate innovations from scientific literature using Large Language Models cites this paper.

Towards unearthing neglected climate innovations from scientific literature using Large Language Models Calibrating LLM-Based Evaluator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:05:54.291216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:05:54.291216Z digest=sha256:39c0d0904263ec63b0dd095e7552ebba3d9b94b4fad4f7004029db9573b9fa07

Observation a57f28fd-7676-4364-9c5f-b975c931315e · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:5248d1d7fedb3461e0e46028d197282afc697ea347b4f85dbfec2e4ef705f394

Observation ff2cbbb3-7af3-4f8b-927d-6c210e994160 · inbound

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator cites this paper.

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Calibrating LLM-Based Evaluator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:15:44.970534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:15:44.970534Z digest=sha256:c19acdd9646dc8e39b3ff2625f7c334d018dc39b6bd1303fad87cf839ea5414a

Observation 4ecfa734-88d1-454a-9f7f-1d66f2b609a9 · inbound

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions cites this paper.

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions Calibrating LLM-Based Evaluator

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-11T20:37:54.919951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:37:54.919951Z digest=sha256:24c66274c688631b5f126a48af617c6d23997e23e36022f1c41bfdbae02c5b36

Observation d7b9ddb2-75e4-42d3-903a-7ffacf660847 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Calibrating LLM-Based Evaluator

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.331051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:40324d179d877320bab2d4eb0b66ecf0c1b2fa10ef3275fe60224156f616390a

Observation 90b2acb0-4145-43dc-98de-32ca74b6b5e4 · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations Calibrating LLM-Based Evaluator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.260875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.260875Z digest=sha256:10a209fa7e2a249a74030c1174c81f103809d14fe7e85b16a9c5267c2e7c715d

Observation 8a8ceceb-ed1b-4e88-874f-d08af2c85220 · inbound

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents cites this paper.

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents Calibrating LLM-Based Evaluator

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.894987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T09:36:50.323723Z digest=sha256:d8f6756c33d2281bbb4298d6565fb55c48d1a72e430e3aae50c5655bcb4f0cf1

Observation 5d5550e5-0e70-4a49-92f3-1286aa587f9c · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Calibrating LLM-Based Evaluator

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:55.567886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:55.567886Z digest=sha256:a1e38dfdc676a14286932773e9e4a307ffe8388e3bdfa7133a177d04bf8ead2b

Observation 3e8a5565-b587-45dd-90df-fa8a37b56243 · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Calibrating LLM-Based Evaluator

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.954717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:50e3c1df8dd12e54c1e35f1c66933d214229e3a1a78358a3ea4050e3b43c0880

Observation 3ae689bf-abfe-4479-937f-c253e9ab8569 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Calibrating LLM-Based Evaluator

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.496695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:54d570b2ef6cbf2eea2982b471b963c9ac67dcb18fdcf90dc1ba7e5746a2f227