Pith. sign in

Paper Citation Record · LEDGER

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.10403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10403 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:13.658964Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T15:14:05.529138Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ef315790-9b74-4f37-80c1-6a150431efaa · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.584381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.584381Z digest=sha256:089f150637313545b1abdfcddaa60f0396a4064e080bd79a6a9d939eddf2b79a

Observation 6acd8fde-133d-4d07-9f32-eedf388fa975 · outbound

This paper cites N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.; 5 Gonzalez, J.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.; 5 Gonzalez, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.882859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:08.700690Z digest=sha256:01b3d990b6fb915b47c2a3045b74c723b9dc233ca9cab0233354ea1ec2aab81d

Observation 113822d6-e3d9-4b00-a9c9-05185fb377f3 · outbound

This paper cites Advances in Neural Information Processing Systems 2023, 36, 46595–46623.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Advances in Neural Information Processing Systems 2023, 36, 46595–46623

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.658458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:08.819739Z digest=sha256:f82c6f4848c913531f5e2cca75a8511d73f3dd497938bfcda6251a9ac41fe94f

Observation a029ec1d-50a9-4bef-9736-cd09ca5b1c94 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Agent-as-a-Judge: Evaluate Agents with Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.958962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.958962Z digest=sha256:da15e7283e3f03ae0852f00373d151a0e446c35aa121ee9cca4edc7841f62596

Observation 290c994f-6f21-4700-9db7-cf29e9336f1d · outbound

This paper cites QuRating: Selecting High-Quality Data for Training Language Models.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation QuRating: Selecting High-Quality Data for Training Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.115749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.115749Z digest=sha256:aa9f35e0fccf7f7a5ecc13244ba7d2715a0b1d3bfb7d185e63ae1a9016032489

Observation a2b0b47b-a7e8-4a48-85b5-8d71ab66010a · outbound

This paper cites F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.485017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:09.307806Z digest=sha256:4ac519660d3f9026b8530b29ca9be14626a2be3da0208a4e9cd60f6a103e34a3

Observation a74c5cd5-db0d-4001-9307-d494038856ad · outbound

This paper cites an unresolved cited work.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:18.243036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:09.434801Z digest=sha256:f16b3975c0df10c99313fb27492c66a4dc6cee9351ac52fb713377d3845c8c35

Observation f980c253-bbe6-46ce-83bc-1e9c8ad4a062 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.627591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.627591Z digest=sha256:6642d79131c934ec670ef9c24bc02fde7dc2d348501a411aa4dd7143bdd23a61

Observation 5d5285b6-e5ed-4824-bb79-f08bef70f2f6 · outbound

This paper cites xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.757567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.757567Z digest=sha256:7a5b7c8567afb96336be9629c19cbab23dc85e57f805f1a4723b8735965c0c62

Observation 19dac899-d440-48c8-8b57-7f5e6991fc71 · outbound

This paper cites Discovering Bias in Latent Space: An Unsupervised Debiasing Approach.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Discovering Bias in Latent Space: An Unsupervised Debiasing Approach

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.900260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.900260Z digest=sha256:e53bf9dd25808e0383263e68ac2b606f85b64bca90a52c8ef05717e0175f94b1

Observation d5d4530c-185d-4291-8809-faa66fc8ed31 · outbound

This paper cites H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.013754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:10.023607Z digest=sha256:58c4aa787a4df0abfdc8934f86504b274c0dfdd6b193556861f52c0375db4026

Observation 8f607aac-142f-4e00-b349-0551a68d1b97 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.196137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:10.196137Z digest=sha256:4567a99d1126aa9d1453a12f546b0cf6bfdf2ca3ba513185e584d22dc243a16f

Observation c07bf243-e9b2-49b9-9f36-fdb29e17b4f4 · outbound

This paper cites J.; De Sa, C.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation J.; De Sa, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.805382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:10.307887Z digest=sha256:1695ac4867bcdf144ac9e1458eddbc82af69c6ba0ebe6e4b93b8d97983cdb4a0

Observation e8c7f540-ff3c-4302-b88c-0909217293e3 · outbound

This paper cites H.; Ehrenberg, H.; Fries, J.; Wu, S.; Ré, C.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation H.; Ehrenberg, H.; Fries, J.; Wu, S.; Ré, C

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.526067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:10.487327Z digest=sha256:85c34b9044c4089f39321851b5816482b44564bbe015c5bb3a7950e8e81b2349

Observation ca798256-8249-4e82-8760-09a04738b7a4 · outbound

This paper cites Training complex models with multi-task weak supervision.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Training complex models with multi-task weak supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.359182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:10.653766Z digest=sha256:dea5498543e975fe5d9f9ddb841f81a9ba9930999476b40e9d8baaf57107de84

Observation affb8cf4-53ab-49ed-bdea-b30ec53d5f5a · outbound

This paper cites Fast and three-rious: Speeding up weak supervision with triplet methods.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Fast and three-rious: Speeding up weak supervision with triplet methods

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.089083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:10.819345Z digest=sha256:8ee907a242dd95789ece85078fb365cbcdb13e67b51ee2e3a0f96d77941e666d

Observation d07bcf6c-d883-4dbd-83f3-8cbd8345e7f8 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation RewardBench: Evaluating Reward Models for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.968389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:10.968389Z digest=sha256:b961a9396bd06502e3f9d857429befe88f67c0e81daeae6f5f9534be70b29f07

Observation 0e7e9cc8-1810-4351-af7f-ae67533c8332 · outbound

This paper cites The Twelfth International Conference on Learning Representations.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation The Twelfth International Conference on Learning Representations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.956616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:11.147019Z digest=sha256:d2df2e42b01d0f56dc38db0818192c0496d4a5e77b973106ba4c88be960ea583

Observation 08063b97-5a06-43dc-b46d-d612811d8fcc · outbound

This paper cites GPT-4o System Card.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.322747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:11.322747Z digest=sha256:3fb03fdc8cc3adb9622a40cfeef00decc6497dc687572c634a813ad702e9c4ea

Observation 7849f027-4f10-4708-9357-5ee5d7c19757 · outbound

This paper cites an unresolved cited work.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:16.873686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:11.471540Z digest=sha256:c70a1f5095308aa19f82202162048cdd5bd41dc0059b15404aadbc25f7b8dc46

Observation 292117a4-594a-45ad-826b-945007f1d7f5 · outbound

This paper cites P.; Fishburne Jr, R.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation P.; Fishburne Jr, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.587401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:11.632589Z digest=sha256:f60606a19035e6b386459b7afe30e28e1a7b060853c36abc46ca606c58dadc7e

Observation 89546b02-b98c-41dd-a339-c03f10bba31b · outbound

This paper cites FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.299829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:11.812496Z digest=sha256:155a73f91c8420c1632ef397f6ce302ee3494b0d139540cf02b7072adfbf1de5

Observation 8562dfc4-bc17-4ef9-bcd6-9bcc0311d637 · outbound

This paper cites Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.018586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:11.939726Z digest=sha256:c5a75e63a7f348cde18db5cba401693a1acbaec2fece12bc393264bcdc5a4af0

Observation 6bdd328f-b03a-4208-8ac1-24a3eca19819 · outbound

This paper cites Universalizing weak supervision.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Universalizing weak supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.672364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:12.071052Z digest=sha256:957ba1d475c88738edd19b1013095de44b9abe98e111fa561779c115170542aa

Observation 5711fb54-8843-4ca0-8982-a659fa51253a · outbound

This paper cites Weak-to-Strong Generalization Through the Data-Centric Lens.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Weak-to-Strong Generalization Through the Data-Centric Lens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.353380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:12.237890Z digest=sha256:c3f69c0c3464a7a8daba31fde550f10a247dc0dea2d4233d765019380af85dd0

Observation f6833448-268e-4a5a-8a56-f5b92c2d2589 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.060847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:12.346391Z digest=sha256:916c804a94314c42ffc07cb4b02188dbae0b161c8d0c43db78e442143900f1d1

Observation bb382b29-30e9-41dc-8491-3df9c32420b7 · outbound

This paper cites GPT-4 Technical Report.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.492714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.492714Z digest=sha256:20c68db9b3cb335201e7133985944ff5282ff70f664cc84e3d89fecbbacabdc2

Observation 1a7646b5-cff3-46c1-b0d7-07c9a51d592e · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Gemma 2: Improving Open Language Models at a Practical Size

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.651746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.651746Z digest=sha256:a7dfa16e3724be13f1c504d72c98e9fe82bae7a6efe4a3d58d1cde8eefc9a0b2

Observation 733cb78b-219d-4772-a6d9-9a4b30d992f6 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.759896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.759896Z digest=sha256:5ddd146ccadcac9375758661b09f1c2330e7e3328385d3e3f5e45805fe4b968b

Observation 86e8dac7-4bd3-4ca3-bfa5-b16cfe8053b8 · outbound

This paper cites Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.934810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.934810Z digest=sha256:7f2aa828bbddbb8d7755c5f8942afe83dab14f93cb73700747bfeb145860d64f

Observation 4c2ad085-1fa2-444a-9242-9cd64b26d309 · outbound

This paper cites The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data Annotators.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data Annotators

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.817946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:13.041107Z digest=sha256:621ca14a63c75c7fbc43e7a93909dbd0fede0cab6d0e260da16a6189c90645b4

Observation abc0dae9-433c-4b1f-96ed-5fa5cfe702e3 · outbound

This paper cites ScriptoriumWS: A Code Generation Assistant for Weak Supervision.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation ScriptoriumWS: A Code Generation Assistant for Weak Supervision

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:34:13.922531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:13.235686Z digest=sha256:832df42a14ff71d4997f3d585190935baa39c20a188e53b4195e0afa067033f6

Observation 5e2df8e1-3ce8-4888-8d33-589c4ba98afa · outbound

This paper cites Autows-bench-101: Benchmarking automated weak supervision with 100 labels.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Autows-bench-101: Benchmarking automated weak supervision with 100 labels

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.616136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:13.376326Z digest=sha256:ff6285340ee61b4553b676cfb3fc74f5d8435092a2de406250e735253e02f3af

Observation cc286cef-bb27-4d85-80b2-35448f585fb9 · outbound

This paper cites Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:13.508340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:13.508340Z digest=sha256:9f4a4b3816229f57cae4bfd602434a620acc777b9f26aa9752f86e7d8e20f1fc

Observation 997eca7d-7afe-4568-bec9-8d04b8b2df7c · outbound

This paper cites ""Calculate readability metrics for response.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation ""Calculate readability metrics for response

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.331717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:34:13.658964Z digest=sha256:70ead90af3f06e76ee1e45756a6c869d4fcd8a866afe67b46525a73f50a12cd3

Pith citing papers

Observation 0f017e17-0894-4340-9915-1d2ef361703d · inbound

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation cites this paper.

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:50.709981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T20:02:40.980538Z digest=sha256:c092650a341932220d07be6df28ababe37287c6f90391226897e3d8c47273617

Observation 286f4207-5d0a-40c4-b6ab-05c718a80828 · inbound

Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation cites this paper.

Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T19:15:44.221307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:15:37.744461Z digest=sha256:6b1e9b811bc6ed1246aac2830c80a2a898ab782be53263d5187d0ff92523f32f

Observation b7d92186-e247-4ba5-ae91-d64872b60041 · inbound

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning cites this paper.

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T15:14:05.529138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T15:14:05.529138Z digest=sha256:0d13c1cadf219756a0b0a491eae741b731d6c8ed5ca5ab34e7278ef007767e58