Pith. sign in

Paper Citation Record · LEDGER

Engineering AI Judge Systems

As of 14 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2411.17793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17793 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:59:18.434546Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:32:47.756343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:33:18.382028Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e07a633d-a329-4f9a-822d-47e8a45c6cad · outbound

This paper cites Claude 3 sonnet has become very lazy,.

Engineering AI Judge Systems Claude 3 sonnet has become very lazy,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.294465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.139088Z digest=sha256:7a0f0a2cdd0e220026d5998daca20430c2bfeb1049957c79436be84f570fb6cc

Observation 0641d58c-b4d9-45ae-a544-f469463b3b76 · outbound

This paper cites Do you guy think the cost of gpt-4 is high,.

Engineering AI Judge Systems Do you guy think the cost of gpt-4 is high,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.271435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.148769Z digest=sha256:c4843f16e21a23d810dc543d9d5628c85f5816b288bfcc369e65a1365b69ebf9

Observation 7e88b40b-4104-444c-a7f7-955aa0354452 · outbound

This paper cites Gpt-4 is crazy expensive,.

Engineering AI Judge Systems Gpt-4 is crazy expensive,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.259355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.152711Z digest=sha256:789db54b288f4bab227e17810a324aea09e70786647e9a6f89cf9a8fc9e49b93

Observation 6e7e3ff0-5b01-49c5-8426-0c64801f4e06 · outbound

This paper cites Open llm leaderboard,.

Engineering AI Judge Systems Open llm leaderboard,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.157331Z digest=sha256:3fefc8fabf169dcfb1e12ad89f2538edfa5ea0c6658913c1ffa1821679112c93

Observation a62f18e3-1362-46cb-9a8a-d5e82ba916dd · outbound

This paper cites Use agent metrics & llm judges to evaluate app performance,.

Engineering AI Judge Systems Use agent metrics & llm judges to evaluate app performance,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.236198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.161619Z digest=sha256:9a9f92272acdbaa95a506b28d7de1ac9e836ea7a1a6e2d03fdbb3778868d95c2

Observation 2d24a91c-5b70-4273-bff3-90967b30d88b · outbound

This paper cites 2030 software engineering,.

Engineering AI Judge Systems 2030 software engineering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.224133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.166669Z digest=sha256:89c4cd4f579b200964fecd5bdea42ed719ffcbf46c58ccfdd297618bd8d85b50

Observation 947c5aeb-8423-4bc7-9dda-6df89c092418 · outbound

This paper cites The acm international conference on the foundations of software en- gineering (fse) 2024,.

Engineering AI Judge Systems The acm international conference on the foundations of software en- gineering (fse) 2024,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.210844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.170796Z digest=sha256:1523b58d231f6932d1846e6efc7353b8cbacb5b0cb184b4e8a72c00a9044b021

Observation 05ca5043-7480-4d7e-8b16-34110261baea · outbound

This paper cites Fm+se summit 2024,.

Engineering AI Judge Systems Fm+se summit 2024,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.198220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.175887Z digest=sha256:a4727727be1af3025c99333b54f9585479bcec580f9c9051068eefea6c85ec27

Observation 40e1b4f1-a253-4d29-a173-96e35aede6e5 · outbound

This paper cites Opea initiative (open platform for enterprise ai (opea),.

Engineering AI Judge Systems Opea initiative (open platform for enterprise ai (opea),

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.186229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.180239Z digest=sha256:b999719d012f53c583cbc4e38ce15d8dc5bd4584ed5440ec82207e6cd244b94f

Observation c3f25090-1dee-424a-9f4d-ef4075c739ff · outbound

This paper cites Re-thinking data strategy and integration for artificial intelligence: concepts, opportuni- ties, and challenges,.

Engineering AI Judge Systems Re-thinking data strategy and integration for artificial intelligence: concepts, opportuni- ties, and challenges,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.172791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.185399Z digest=sha256:6ee49d42d5ccbe048a449b289e533bafde281d92f74bffdaa97633178d1792d5

Observation 8a957eb4-3f66-4202-b0f1-8923229c292d · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Engineering AI Judge Systems Constitutional AI: Harmlessness from AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.189110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.189110Z digest=sha256:6e6ed63adf29a9e8291ddd2a0528d41fdda2dfbe22b96ed60d36e5eaf18b660b

Observation c8023431-371b-4128-96f9-6d4dd42ca5d2 · outbound

This paper cites Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs.

Engineering AI Judge Systems Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.194285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.194285Z digest=sha256:bc917ffa6c82387620b5a9e1792fcfc0d8fd92e3ffcb57b7219d1165980c21c6

Observation 92cf6acd-09fd-4426-b0ef-6bf32fe48823 · outbound

This paper cites Meteor: an automatic metric for MT evalua- tion with improved correlation with human judgments,.

Engineering AI Judge Systems Meteor: an automatic metric for MT evalua- tion with improved correlation with human judgments,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.159428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.198230Z digest=sha256:17aeedd1a49dfb220256c75660c4b4913de5fdb88495b0b2bd332cb8014b1aa3

Observation cfb058a0-0668-4acf-87dd-305edac4a074 · outbound

This paper cites Test driven development: By example addison-wesley,.

Engineering AI Judge Systems Test driven development: By example addison-wesley,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.147267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.201903Z digest=sha256:7331b1fa01eabd9c93cd0294f557f4dc911455a3adc42b035549037b65afb8f6

Observation 71493e67-f94e-479d-a0cd-602bac9facd1 · outbound

This paper cites On the dangers of stochastic parrots: can language models be too big?.

Engineering AI Judge Systems On the dangers of stochastic parrots: can language models be too big?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.135774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.205397Z digest=sha256:1fb39b3874f7237b68a7d226a7861710d19513bbacacbecf5d368a647cb29659

Observation c2ac4a3a-4cff-48fd-97bf-e8fd9f8a5dd5 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Engineering AI Judge Systems Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.208969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.208969Z digest=sha256:6298464339fee9d2d990f8a79f9083adf62cdf9364a868ed50b358add1f3443e

Observation 62e3339c-7943-452c-847f-8cbde46d5111 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Engineering AI Judge Systems ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.212705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.212705Z digest=sha256:a7633946ca559e94d8640717543f24283b79d784d751bfbb689d033d8d7d51ec

Observation 0935c594-8e41-416a-95db-9f86a061e380 · outbound

This paper cites A survey on evaluation of large language models,.

Engineering AI Judge Systems A survey on evaluation of large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.122927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.217544Z digest=sha256:1635956b77258ca52d7487da4a4ed7440386986cece2eef791ce50309c8269ea

Observation e65b98a8-a1b8-4100-a638-188fd934f478 · outbound

This paper cites Unleashing the potential of prompt engineering for large language models.

Engineering AI Judge Systems Unleashing the potential of prompt engineering for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.221024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.221024Z digest=sha256:9ce9c8b4f211719c7f8567b4a4ebd45964b15643212fbe867546f041eba4d85b

Observation a970bc9c-49d3-465b-9583-2f2c3cd9b86e · outbound

This paper cites Towards training reproducible deep learning models,.

Engineering AI Judge Systems Towards training reproducible deep learning models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.110698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.224849Z digest=sha256:38672d5bfe834148e12168395f046a8c029aaa66e826c6a1e714ad442a9fb0b7

Observation d2fa95c7-8c31-498e-8bdf-dec0558784c8 · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

Engineering AI Judge Systems Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.228507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.228507Z digest=sha256:8ba47a1142510d12b1ce1317d68c68f4512408f23d4f45cb07272ad07933c831

Observation c5ad56e2-3a5c-4c91-a84c-08c4ee58a9c2 · outbound

This paper cites How is ChatGPT's behavior changing over time?.

Engineering AI Judge Systems How is ChatGPT's behavior changing over time?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.233491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.233491Z digest=sha256:f803da1bb2f35801e9d8b8d6e44ef59dfb20ee3c8b7e8097a80df1a3c2055150

Observation 64e25685-2f9e-4c1b-9113-cc8e95572a23 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Engineering AI Judge Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.237211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.237211Z digest=sha256:417ebb94f6ac17490a0a0caa391e2cef515846162db50da5a2926cf3da507050

Observation 2ddc37fd-3306-4aac-bcb2-a6cc389bc29f · outbound

This paper cites Available: https://www.reddit.com/r/ClaudeAI/comments/ 1bv8ww5/claude 3 sonnet has become very lazy/.

Engineering AI Judge Systems Available: https://www.reddit.com/r/ClaudeAI/comments/ 1bv8ww5/claude 3 sonnet has become very lazy/

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.282971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.143956Z digest=sha256:014ec19e22aed868e11cb15c4d712ff6de327628490a9811bab9dcf95f478d65

Observation 10d1d72c-38ce-4f89-8809-379d44341fc4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Engineering AI Judge Systems Training Verifiers to Solve Math Word Problems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.241194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.241194Z digest=sha256:7c6b5b51ab826071825f2e2bfc12099f0201dec0ab728be1c180e6ee485f8fe2

Observation 54a73f2e-70a0-43f2-89e7-684da527c5e6 · outbound

This paper cites Evalullm: llm assisted evaluation of generative outputs,.

Engineering AI Judge Systems Evalullm: llm assisted evaluation of generative outputs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.098115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.245859Z digest=sha256:303d2df34d77f074c59425865f65c8e1c2318debd88cce4a555f37528cc5f2cf

Observation 2ab6465c-c725-4a04-b939-394da8855761 · outbound

This paper cites Qlora: ef- ficient finetuning of quantized llms,.

Engineering AI Judge Systems Qlora: ef- ficient finetuning of quantized llms,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.086396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.249801Z digest=sha256:ad856771f03bbd900700e2217c09d352cd23fe09bd323309fff5e7d94ea15b6a

Observation aec2f2f2-46ed-4a8e-9983-d24f1b299583 · outbound

This paper cites Fira: fine-grained graph-based code change representation for automated com- mit message generation,.

Engineering AI Judge Systems Fira: fine-grained graph-based code change representation for automated com- mit message generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.075156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.253234Z digest=sha256:b34c9f2d15b246a1abf26fc83def287d9d0b22aa1235def40d20043e1062b2b8

Observation 2c4f6a65-7ce3-4ec6-a895-0cbd0da38f37 · outbound

This paper cites Alpacafarm: a simulation framework for methods that learn from human feedback,.

Engineering AI Judge Systems Alpacafarm: a simulation framework for methods that learn from human feedback,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.063807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.256879Z digest=sha256:f0badcedf4bf80264dac74dbcd559c2d46bf937381b1b14ea8a38345423e48e6

Observation 9831402a-94fc-4b4a-bb73-8d52cc27b3c7 · outbound

This paper cites Bias and fairness in large language models: a survey,.

Engineering AI Judge Systems Bias and fairness in large language models: a survey,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.052740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.260862Z digest=sha256:7299577ab1751adb3c2d2ca207e6d62eb5249bcf78dc0ba2483653b37ad02d29

Observation 535c660a-2eb5-4e4a-8c8c-2d803839f901 · outbound

This paper cites Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning.

Engineering AI Judge Systems Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.266978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.266978Z digest=sha256:d62c928d92c3e4c35e70c1920ee7ccd9e9fdbb32819de2806036d2d48c7f758d

Observation 4944e0f9-35d0-4078-8829-e8f5a610a8e9 · outbound

This paper cites LLM-based NLG Evaluation: Current Status and Challenges.

Engineering AI Judge Systems LLM-based NLG Evaluation: Current Status and Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.271869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.271869Z digest=sha256:ec1c5cc6e3b2b004744bec33ff56d80f8973971be0127f70c6a41061a6904ebf

Observation 964d0df2-5403-4e1b-8136-a64903194b90 · outbound

This paper cites Fm+se vision 2030,.

Engineering AI Judge Systems Fm+se vision 2030,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.041370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.276954Z digest=sha256:6c998410067bb2b4993d1214996ba51069bc4630f7f9816b4a257baaed434211

Observation 80f2e391-f7a0-4f0b-9247-39c261d7d047 · outbound

This paper cites Rethinking software engineering in the foundation model era: a curated catalogue of challenges in the development of trustworthy fmware,.

Engineering AI Judge Systems Rethinking software engineering in the foundation model era: a curated catalogue of challenges in the development of trustworthy fmware,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.029239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.280771Z digest=sha256:a426f33c78d56bb1173b2360e73b7bbecfc01548289ca2036128798221abea26

Observation 38c394b4-fb5d-4586-b09d-45b121f5a5ef · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Engineering AI Judge Systems Measuring Massive Multitask Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.285054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.285054Z digest=sha256:8d31d2d0e7fb619fe95e0f7a19ca7724f8b9ed25809e9a365934ee43f6bde2e5

Observation abc6d780-04f7-4f85-bcb2-8b0d8f8b217b · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Engineering AI Judge Systems A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.289124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.289124Z digest=sha256:a160a8c994efc1d051877e71de67ddeffa8c9f4c31ae49d4f0589e3b87808415

Observation 8f6c5c9f-b355-4868-9765-adf08a204dd1 · outbound

This paper cites AI safety via debate.

Engineering AI Judge Systems AI safety via debate

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.293117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.293117Z digest=sha256:de5a63d423bd92e4d723630f6f35343bd8d7863927499c936858fcb05e617cc0

Observation 601eb689-23ad-4590-b3ee-368c505574dd · outbound

This paper cites Kejriwal, Domain-specific knowledge graph construction.

Engineering AI Judge Systems Kejriwal, Domain-specific knowledge graph construction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.018229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.297827Z digest=sha256:f76c06b75e23d6959b79f2a80530bfcaac515799378f66bfe82afd016e6acc85

Observation dce9d1bd-1dd9-479a-9f89-db458e0f2469 · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

Engineering AI Judge Systems On scalable oversight with weak LLMs judging strong LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.301527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.301527Z digest=sha256:efc4ff050770fb5b6aef3b835cbb1c534ea2d6847423dc41251cc2ff6bdf2e39

Observation 28284a8d-410a-4e89-b44b-1d1c75aff2f3 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Engineering AI Judge Systems Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.305344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.305344Z digest=sha256:b103c9b735f481d206d58d4140f9e4a21dfc4c6d4f0dfd87793a57991ee0006a

Observation 84a337a4-3562-4833-9efb-ddcd9b7d00f5 · outbound

This paper cites Dspy: compiling declarative language model calls into state-of-the-art pipelines,.

Engineering AI Judge Systems Dspy: compiling declarative language model calls into state-of-the-art pipelines,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.007287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.309227Z digest=sha256:9748a03a027583c0c3775c07ea6964bb9fb2d58f6eb07c3511e180b90cd2a3dd

Observation f9386766-e1fd-461f-9013-e120f2ed1260 · outbound

This paper cites Software engineering for machine learning applications (semla) 2023,.

Engineering AI Judge Systems Software engineering for machine learning applications (semla) 2023,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.993839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.313972Z digest=sha256:9236cece0b52e60c338a1dbb674fa8d3372366d5affd517e33f0389fdf3b5116

Observation dac9e283-2527-47fb-8842-c2df3acbb6cc · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Engineering AI Judge Systems Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.317823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.317823Z digest=sha256:c3fbd6e910da081a59b08ed7371dceb125205deca236793dcecbab0ed748b721

Observation 3aec817e-5fca-4029-af5c-2709ab851d84 · outbound

This paper cites On the role of knowledge graphs in explainable ai,.

Engineering AI Judge Systems On the role of knowledge graphs in explainable ai,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.981263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.322382Z digest=sha256:9875341cfdf02f27c0bf38bfaacc72d4ad2fe1ab4ffe678bbb973046d7702b78

Observation 2104f4a8-8d00-479d-aa47-f19387188b8c · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Engineering AI Judge Systems Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.326017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.326017Z digest=sha256:cc79e83a23727af8f721eb936aecad56d205737e66ebf0a424a6d685a50b297a

Observation 9724063b-752b-41cc-92a5-2e8b9e3fa414 · outbound

This paper cites Rouge: a package for automatic evaluation of summaries,.

Engineering AI Judge Systems Rouge: a package for automatic evaluation of summaries,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.968535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.330973Z digest=sha256:b3a041e5f393c0857e75c6f7dd0d403e1eba3d2d8f81736b8d072aeea9d4604c

Observation 56a5924d-8edc-4c5f-8af5-dc72367b80be · outbound

This paper cites Best Practices and Lessons Learned on Synthetic Data.

Engineering AI Judge Systems Best Practices and Lessons Learned on Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.334631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.334631Z digest=sha256:d351016afb503db97bbf37afccb4b832ba8fd23d573c471c986c3d9b623947ef

Observation a57f28fd-7676-4364-9c5f-b975c931315e · outbound

This paper cites Calibrating LLM-Based Evaluator.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:89ee2d1530c0a5326e450974ed6095a08192855021f7dc740d364a044805e744

Observation ae9b3172-cb8b-4264-b00e-b1fc5795ef68 · outbound

This paper cites Generating training data with language models: towards zero-shot language understanding,.

Engineering AI Judge Systems Generating training data with language models: towards zero-shot language understanding,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.956636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.343429Z digest=sha256:7bdc773d53c7ef3b04cb7d59c03b3ea0f32f6ea16e4b19d1dcdf4e0258c148d6

Observation eb717707-4edb-4c03-aeb5-23bf499ef0a4 · outbound

This paper cites Cider: robust consensus-based image description evaluation,.

Engineering AI Judge Systems Cider: robust consensus-based image description evaluation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.944650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.348047Z digest=sha256:dfa647fe84020384469b0ae82171663a0b8e7ccbc65765c9509b362655e91d26

Observation 95a55474-bc6c-4ec1-81f6-73192c62e639 · outbound

This paper cites A framework for evaluating and improving requirements specifications based on the developers and testers perspective,.

Engineering AI Judge Systems A framework for evaluating and improving requirements specifications based on the developers and testers perspective,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.932435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.352122Z digest=sha256:c197cf765a8a134348472e099906c536a630a7b80a15d6bd36b2154dcf135639

Observation 698c9c48-3d7c-4663-9190-028b2013d533 · outbound

This paper cites Human-Centered Design Recommendations for LLM-as-a-Judge.

Engineering AI Judge Systems Human-Centered Design Recommendations for LLM-as-a-Judge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.356623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.356623Z digest=sha256:a3b778669195ebd13b525240f216a7ffee9271351b01f0ff742c880f5103c1b7

Observation d380be56-886c-43af-a80c-d04386398aaf · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Engineering AI Judge Systems Bleu: a method for automatic evaluation of machine translation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.920316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.360737Z digest=sha256:ec8cfde74b6f039bbe44e6040648b45792f46fb87fecb839a84011a9722c24bc

Observation a5a90d46-9e76-4ecb-94ce-4bf898a63156 · outbound

This paper cites The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities.

Engineering AI Judge Systems The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.365781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.365781Z digest=sha256:3a931c40a00932bca4cc13810c0e8b1ce72403e411aa8ab6a2748257b65ff5c8

Observation 57920572-3461-49cf-b1b9-5a97faa05cee · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Engineering AI Judge Systems Verbosity Bias in Preference Labeling by Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.369902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.369902Z digest=sha256:64a3e48853f853f0f4a564aee42fbfb7f14b0ea51009336d95c1c93d1f63ef81

Observation 0f1b10a0-d706-48ae-a6fd-e0e483db6f0e · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

Engineering AI Judge Systems BLEURT: Learning Robust Metrics for Text Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.374841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.374841Z digest=sha256:8c7b13e46dd4a7846763b72892cb8d6c4f53f206e5bde979962615cbc7fe5318

Observation 1cd8717b-f6b1-4089-aec4-d5b796dbcae0 · outbound

This paper cites On automatic summarization of what and why information in source code changes,.

Engineering AI Judge Systems On automatic summarization of what and why information in source code changes,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.906996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.379788Z digest=sha256:f5fd6b0255d6c15fb4e05d4865044ebd54ab267674f64c20cf42a146275e974c

Observation ff1d24d0-ab0f-4ea3-92d4-57d3d1841f88 · outbound

This paper cites On the evaluation of neural code summarization,.

Engineering AI Judge Systems On the evaluation of neural code summarization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.894287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.383721Z digest=sha256:653f18a447cf3ca9a1e48ba384ff4df4c235c57048ece48f60e0d66b6edb473b

Observation f1f2be3f-b8ef-4e6c-9a72-9c183b1bb0f9 · outbound

This paper cites RACE: Retrieval-augmented commit message generation,.

Engineering AI Judge Systems RACE: Retrieval-augmented commit message generation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.882388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.388673Z digest=sha256:a0979a3acb77be8a6bb32afafe3030d5137b91a3e1680f1909c609505f6c408e

Observation 9de57533-72ab-447b-9c2e-9a16d1936911 · outbound

This paper cites On the evaluation of commit message generation models: an experimental study,.

Engineering AI Judge Systems On the evaluation of commit message generation models: an experimental study,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.870426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.392463Z digest=sha256:88e8137c36aada483266d2e9bbea1c0fedfe8057f854b451584ff059d22d2878

Observation 68bbccb3-869f-49bb-aa65-0f68b69ef797 · outbound

This paper cites Alpaca: a strong, replicable instruction-following model,.

Engineering AI Judge Systems Alpaca: a strong, replicable instruction-following model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.858724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.397190Z digest=sha256:e2a74baaba3fde02d33fb519348fc101e5dd1ea2eada6ba812862d81f70988f6

Observation 196eee81-2147-4992-9d96-4f60defcddda · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Engineering AI Judge Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.400788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.400788Z digest=sha256:97d225d32629d3b44a5f50e87f2c338a270e713c025aaa3c8c80a5c88224db66

Observation d64aa7a2-c591-4f8b-8884-5cb6100a9cc3 · outbound

This paper cites Synthetic data, real errors: how (not) to publish and use synthetic data,.

Engineering AI Judge Systems Synthetic data, real errors: how (not) to publish and use synthetic data,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.844266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.405510Z digest=sha256:bde123ae924c43eacfd45278366a15ef58a5eb7a2e2fb34cc0b7b96129e46fdc

Observation a7d2c3be-6852-4eea-b3be-9e62162e1d76 · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Engineering AI Judge Systems Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.409126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.409126Z digest=sha256:0ab3ff1ef5f12e86e60741ad8b59bc6a9e2480b269d55df806de7b3744d8acc9

Observation 02d0b65c-f5f1-4087-b51f-22bad0f786a2 · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Engineering AI Judge Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.413886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.413886Z digest=sha256:d3d3510595f7fc7ccf2cb67a438e04c94611cb42aff60abb2f0733c65c829fdd

Observation 5746f32f-31dc-40ec-b425-a3aa15be497f · outbound

This paper cites Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems.

Engineering AI Judge Systems Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.417892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.417892Z digest=sha256:51bc2a42d0fb250afe5b9dfde947a90c406bdf624e8137e4795a0ae04d364094

Observation 99eedc02-be11-4d1c-bb22-864eaef6e7e4 · outbound

This paper cites Fake it till you make it: face analysis in the wild using synthetic data alone,.

Engineering AI Judge Systems Fake it till you make it: face analysis in the wild using synthetic data alone,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.831987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.422862Z digest=sha256:e8704179367e5a96a5e473a7e150aacfd7fa74faace87b4e39e34e1afbb51eec

Observation 219aebda-04d4-455d-b19a-6145c4cbea0a · outbound

This paper cites Commit message generation via chatgpt: how far are we?.

Engineering AI Judge Systems Commit message generation via chatgpt: how far are we?

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.820449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.427237Z digest=sha256:808ce5d87f3563e4bf6eed05ab0c6c9082b6f4f14ba5501358a78f107b07b69b

Observation f890a40e-3532-4f3d-b0f0-0e99dbe9cb2c · outbound

This paper cites Automatic commit message generation: a critical review and directions for future work,.

Engineering AI Judge Systems Automatic commit message generation: a critical review and directions for future work,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.808693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.431034Z digest=sha256:365d4fd5cf2a68a12fd4c1531eaa63bd60f94962fb284bce38a0afc1277f2b82

Observation b3e532f0-e1c9-45c0-8e81-6851a8f47262 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Engineering AI Judge Systems Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.795757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:59:18.434546Z digest=sha256:cdaca715f71b554d31664d3162b0a87f0536ab6f6383a0a3405371e3a3149594

Pith citing papers

Observation a1f08563-e1d9-450b-b1fb-942ab696aa28 · inbound

Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers cites this paper.

Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers Engineering AI Judge Systems

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:32:15.515399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T09:32:08.596259Z digest=sha256:41583265bfd1918620e9777a95549122203b7f51a224ddd547efcf58339b18c4

Observation c3ec2d89-6bea-4e6b-8f3b-09f152b10195 · inbound

An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges cites this paper.

An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges Engineering AI Judge Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:18.383835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T10:32:47.756343Z digest=sha256:0d1d9e09ee4eef7b83515590cc885c7e54d5f67b16950a2631cade26e8b558de