Pith. sign in

Paper Citation Record · LEDGER

Engineering AI Judge Systems

As of 13 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2411.17793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17793 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:59:18.434546Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:32:47.756343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:33:18.382028Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e07a633d-a329-4f9a-822d-47e8a45c6cad · outbound

This paper cites Claude 3 sonnet has become very lazy,.

Engineering AI Judge Systems Claude 3 sonnet has become very lazy,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.294465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.139088Z digest=sha256:144fe281e862a9216db47c92550114f136b32af99c68312b2aa9098cc0892d3a

Observation 0641d58c-b4d9-45ae-a544-f469463b3b76 · outbound

This paper cites Do you guy think the cost of gpt-4 is high,.

Engineering AI Judge Systems Do you guy think the cost of gpt-4 is high,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.271435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.148769Z digest=sha256:97762e5bbe73edd653fd7578007034d86414aa73ae2746396ab02bbe320b8634

Observation 7e88b40b-4104-444c-a7f7-955aa0354452 · outbound

This paper cites Gpt-4 is crazy expensive,.

Engineering AI Judge Systems Gpt-4 is crazy expensive,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.259355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.152711Z digest=sha256:c00a42b3689410308454578966b37306c94690057cc78bae8ce033cd9c8bf2cc

Observation 6e7e3ff0-5b01-49c5-8426-0c64801f4e06 · outbound

This paper cites Open llm leaderboard,.

Engineering AI Judge Systems Open llm leaderboard,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.157331Z digest=sha256:fb34471942ed071cd8b8aafcf7d3671800c722b837340b96ef1aa13906e3bc4f

Observation a62f18e3-1362-46cb-9a8a-d5e82ba916dd · outbound

This paper cites Use agent metrics & llm judges to evaluate app performance,.

Engineering AI Judge Systems Use agent metrics & llm judges to evaluate app performance,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.236198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.161619Z digest=sha256:c3fd920d36e9b8c34103e858043afcb83a9dcaf5813df9204e6e3ed6b1b28ddf

Observation 2d24a91c-5b70-4273-bff3-90967b30d88b · outbound

This paper cites 2030 software engineering,.

Engineering AI Judge Systems 2030 software engineering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.224133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.166669Z digest=sha256:21e76b17cf8763123fc16c3a5737b83669fbd5996dc968d6608f596a3b624cfb

Observation 947c5aeb-8423-4bc7-9dda-6df89c092418 · outbound

This paper cites The acm international conference on the foundations of software en- gineering (fse) 2024,.

Engineering AI Judge Systems The acm international conference on the foundations of software en- gineering (fse) 2024,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.210844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.170796Z digest=sha256:d74ed379689ac542a5f82007ccb32e51a98472722374ed6ee00ddb2aa4478d05

Observation 05ca5043-7480-4d7e-8b16-34110261baea · outbound

This paper cites Fm+se summit 2024,.

Engineering AI Judge Systems Fm+se summit 2024,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.198220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.175887Z digest=sha256:7582fde0989a105a5db97ed7ac7921a05c9da3292f952587a1d6279b3ac6d4b1

Observation 40e1b4f1-a253-4d29-a173-96e35aede6e5 · outbound

This paper cites Opea initiative (open platform for enterprise ai (opea),.

Engineering AI Judge Systems Opea initiative (open platform for enterprise ai (opea),

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.186229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.180239Z digest=sha256:d12b78e9a77bc0bc064e3f85beca1bd2a43c89e9cb7bca7429b39ce39cf973af

Observation c3f25090-1dee-424a-9f4d-ef4075c739ff · outbound

This paper cites Re-thinking data strategy and integration for artificial intelligence: concepts, opportuni- ties, and challenges,.

Engineering AI Judge Systems Re-thinking data strategy and integration for artificial intelligence: concepts, opportuni- ties, and challenges,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.172791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.185399Z digest=sha256:81d93b96933abcbb7dcac6bf3ce0c2e67af2c8e32e92e5f5bffa6a51234a1e47

Observation 8a957eb4-3f66-4202-b0f1-8923229c292d · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Engineering AI Judge Systems Constitutional AI: Harmlessness from AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.189110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.189110Z digest=sha256:6e6ed63adf29a9e8291ddd2a0528d41fdda2dfbe22b96ed60d36e5eaf18b660b

Observation c8023431-371b-4128-96f9-6d4dd42ca5d2 · outbound

This paper cites Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs.

Engineering AI Judge Systems Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.194285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.194285Z digest=sha256:bc917ffa6c82387620b5a9e1792fcfc0d8fd92e3ffcb57b7219d1165980c21c6

Observation 92cf6acd-09fd-4426-b0ef-6bf32fe48823 · outbound

This paper cites Meteor: an automatic metric for MT evalua- tion with improved correlation with human judgments,.

Engineering AI Judge Systems Meteor: an automatic metric for MT evalua- tion with improved correlation with human judgments,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.159428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.198230Z digest=sha256:3b6513e8d6329e5813c173f142b0b4ca87c5c4f52de4b8278ab4683ff6635d5e

Observation cfb058a0-0668-4acf-87dd-305edac4a074 · outbound

This paper cites Test driven development: By example addison-wesley,.

Engineering AI Judge Systems Test driven development: By example addison-wesley,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.147267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.201903Z digest=sha256:b4fb45a6af61ccf5ab90289ad526270b65b8746220495e8530546321f002399a

Observation 71493e67-f94e-479d-a0cd-602bac9facd1 · outbound

This paper cites On the dangers of stochastic parrots: can language models be too big?.

Engineering AI Judge Systems On the dangers of stochastic parrots: can language models be too big?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.135774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.205397Z digest=sha256:214457593e0c65076ca2dd978d7077b056c44f14eed0285f21ac1c210b206cda

Observation c2ac4a3a-4cff-48fd-97bf-e8fd9f8a5dd5 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Engineering AI Judge Systems Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.208969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.208969Z digest=sha256:6298464339fee9d2d990f8a79f9083adf62cdf9364a868ed50b358add1f3443e

Observation 62e3339c-7943-452c-847f-8cbde46d5111 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Engineering AI Judge Systems ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.212705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.212705Z digest=sha256:a7633946ca559e94d8640717543f24283b79d784d751bfbb689d033d8d7d51ec

Observation 0935c594-8e41-416a-95db-9f86a061e380 · outbound

This paper cites A survey on evaluation of large language models,.

Engineering AI Judge Systems A survey on evaluation of large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.122927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.217544Z digest=sha256:807aa069b984e91e36f8f89e6a3be87aacb5c211e800b715a9335e1d28fc96c9

Observation e65b98a8-a1b8-4100-a638-188fd934f478 · outbound

This paper cites Unleashing the potential of prompt engineering for large language models.

Engineering AI Judge Systems Unleashing the potential of prompt engineering for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.221024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.221024Z digest=sha256:9ce9c8b4f211719c7f8567b4a4ebd45964b15643212fbe867546f041eba4d85b

Observation a970bc9c-49d3-465b-9583-2f2c3cd9b86e · outbound

This paper cites Towards training reproducible deep learning models,.

Engineering AI Judge Systems Towards training reproducible deep learning models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.110698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.224849Z digest=sha256:1967544525c52bf1d49ca4f8a9f496070ef7c4e48c7304b2cef4030d6c17e4ee

Observation d2fa95c7-8c31-498e-8bdf-dec0558784c8 · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

Engineering AI Judge Systems Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.228507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.228507Z digest=sha256:8ba47a1142510d12b1ce1317d68c68f4512408f23d4f45cb07272ad07933c831

Observation c5ad56e2-3a5c-4c91-a84c-08c4ee58a9c2 · outbound

This paper cites How is ChatGPT's behavior changing over time?.

Engineering AI Judge Systems How is ChatGPT's behavior changing over time?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.233491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.233491Z digest=sha256:204b85ac159d5dfcca487d556d0a30f5be5ad362e257c393207e8a4076cd35c5

Observation 64e25685-2f9e-4c1b-9113-cc8e95572a23 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Engineering AI Judge Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.237211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.237211Z digest=sha256:417ebb94f6ac17490a0a0caa391e2cef515846162db50da5a2926cf3da507050

Observation 2ddc37fd-3306-4aac-bcb2-a6cc389bc29f · outbound

This paper cites Available: https://www.reddit.com/r/ClaudeAI/comments/ 1bv8ww5/claude 3 sonnet has become very lazy/.

Engineering AI Judge Systems Available: https://www.reddit.com/r/ClaudeAI/comments/ 1bv8ww5/claude 3 sonnet has become very lazy/

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.282971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.143956Z digest=sha256:551523448e2eef5168ab3a4abbb2988e4d92657a0bbcac71adabd221967addc1

Observation 10d1d72c-38ce-4f89-8809-379d44341fc4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Engineering AI Judge Systems Training Verifiers to Solve Math Word Problems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.241194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.241194Z digest=sha256:8d871ba169f7c1eecda0290431b01f5e29f29f5af6322e4d3fbfba041e9a23f4

Observation 54a73f2e-70a0-43f2-89e7-684da527c5e6 · outbound

This paper cites Evalullm: llm assisted evaluation of generative outputs,.

Engineering AI Judge Systems Evalullm: llm assisted evaluation of generative outputs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.098115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.245859Z digest=sha256:a78197d6023be54083f3555c9c93c0194ea7f42ee9620cf100d1ae8fe762c6a8

Observation 2ab6465c-c725-4a04-b939-394da8855761 · outbound

This paper cites Qlora: ef- ficient finetuning of quantized llms,.

Engineering AI Judge Systems Qlora: ef- ficient finetuning of quantized llms,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.086396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.249801Z digest=sha256:4aa6894551f28e9d96828579e29171b2fb70335b4f77ffc64999df10f06853a8

Observation aec2f2f2-46ed-4a8e-9983-d24f1b299583 · outbound

This paper cites Fira: fine-grained graph-based code change representation for automated com- mit message generation,.

Engineering AI Judge Systems Fira: fine-grained graph-based code change representation for automated com- mit message generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.075156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.253234Z digest=sha256:fa0cfb597bc955ac3f46eb436e3fe3e104c94bc53ab833bd5a78e184196e736f

Observation 2c4f6a65-7ce3-4ec6-a895-0cbd0da38f37 · outbound

This paper cites Alpacafarm: a simulation framework for methods that learn from human feedback,.

Engineering AI Judge Systems Alpacafarm: a simulation framework for methods that learn from human feedback,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.063807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.256879Z digest=sha256:286fc3a5e977e138f986c06c46bdb8541ef034ee5e6c3148d508a52cd5b60798

Observation 9831402a-94fc-4b4a-bb73-8d52cc27b3c7 · outbound

This paper cites Bias and fairness in large language models: a survey,.

Engineering AI Judge Systems Bias and fairness in large language models: a survey,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.052740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.260862Z digest=sha256:482d7bf983ecee417a9bd254f0648305c6ee444806620c2026ba411612df2a4d

Observation 535c660a-2eb5-4e4a-8c8c-2d803839f901 · outbound

This paper cites Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning.

Engineering AI Judge Systems Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.266978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.266978Z digest=sha256:d62c928d92c3e4c35e70c1920ee7ccd9e9fdbb32819de2806036d2d48c7f758d

Observation 4944e0f9-35d0-4078-8829-e8f5a610a8e9 · outbound

This paper cites LLM-based NLG Evaluation: Current Status and Challenges.

Engineering AI Judge Systems LLM-based NLG Evaluation: Current Status and Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.271869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.271869Z digest=sha256:ec1c5cc6e3b2b004744bec33ff56d80f8973971be0127f70c6a41061a6904ebf

Observation 964d0df2-5403-4e1b-8136-a64903194b90 · outbound

This paper cites Fm+se vision 2030,.

Engineering AI Judge Systems Fm+se vision 2030,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.041370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.276954Z digest=sha256:a078768571843bb829b844c757f6a43dc45941f9e52cacc337ac859dc48fc24a

Observation 80f2e391-f7a0-4f0b-9247-39c261d7d047 · outbound

This paper cites Rethinking software engineering in the foundation model era: a curated catalogue of challenges in the development of trustworthy fmware,.

Engineering AI Judge Systems Rethinking software engineering in the foundation model era: a curated catalogue of challenges in the development of trustworthy fmware,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.029239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.280771Z digest=sha256:3189fdc75e00757e8eee13a8815d8daf56f7bb1522941795e0b9d3349a7791e4

Observation 38c394b4-fb5d-4586-b09d-45b121f5a5ef · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Engineering AI Judge Systems Measuring Massive Multitask Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.285054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.285054Z digest=sha256:8d31d2d0e7fb619fe95e0f7a19ca7724f8b9ed25809e9a365934ee43f6bde2e5

Observation abc6d780-04f7-4f85-bcb2-8b0d8f8b217b · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Engineering AI Judge Systems A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.289124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.289124Z digest=sha256:a160a8c994efc1d051877e71de67ddeffa8c9f4c31ae49d4f0589e3b87808415

Observation 8f6c5c9f-b355-4868-9765-adf08a204dd1 · outbound

This paper cites AI safety via debate.

Engineering AI Judge Systems AI safety via debate

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.293117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.293117Z digest=sha256:de5a63d423bd92e4d723630f6f35343bd8d7863927499c936858fcb05e617cc0

Observation 601eb689-23ad-4590-b3ee-368c505574dd · outbound

This paper cites Kejriwal, Domain-specific knowledge graph construction.

Engineering AI Judge Systems Kejriwal, Domain-specific knowledge graph construction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.018229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.297827Z digest=sha256:a9e3e3eec49b9982447dc665f608f7f02a20e9137d79feaa9c4b1b861790a6b0

Observation dce9d1bd-1dd9-479a-9f89-db458e0f2469 · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

Engineering AI Judge Systems On scalable oversight with weak LLMs judging strong LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.301527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.301527Z digest=sha256:efc4ff050770fb5b6aef3b835cbb1c534ea2d6847423dc41251cc2ff6bdf2e39

Observation 28284a8d-410a-4e89-b44b-1d1c75aff2f3 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Engineering AI Judge Systems Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.305344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.305344Z digest=sha256:b103c9b735f481d206d58d4140f9e4a21dfc4c6d4f0dfd87793a57991ee0006a

Observation 84a337a4-3562-4833-9efb-ddcd9b7d00f5 · outbound

This paper cites Dspy: compiling declarative language model calls into state-of-the-art pipelines,.

Engineering AI Judge Systems Dspy: compiling declarative language model calls into state-of-the-art pipelines,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.007287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.309227Z digest=sha256:54a1b8d35674631f541745340e18af1eeeed887eddd5abe29f1e1a8bfb88a4c2

Observation f9386766-e1fd-461f-9013-e120f2ed1260 · outbound

This paper cites Software engineering for machine learning applications (semla) 2023,.

Engineering AI Judge Systems Software engineering for machine learning applications (semla) 2023,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.993839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.313972Z digest=sha256:4d56eef1c06aa9cea5275ef7a45602aaa78ccc4b8b75bb57d07fc31846498c4d

Observation dac9e283-2527-47fb-8842-c2df3acbb6cc · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Engineering AI Judge Systems Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.317823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.317823Z digest=sha256:c3fbd6e910da081a59b08ed7371dceb125205deca236793dcecbab0ed748b721

Observation 3aec817e-5fca-4029-af5c-2709ab851d84 · outbound

This paper cites On the role of knowledge graphs in explainable ai,.

Engineering AI Judge Systems On the role of knowledge graphs in explainable ai,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.981263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.322382Z digest=sha256:db5eaef65a7f38b74257099b40dd02e03c00f67d835b4c649749e0404f562c6e

Observation 2104f4a8-8d00-479d-aa47-f19387188b8c · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Engineering AI Judge Systems Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.326017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.326017Z digest=sha256:cc79e83a23727af8f721eb936aecad56d205737e66ebf0a424a6d685a50b297a

Observation 9724063b-752b-41cc-92a5-2e8b9e3fa414 · outbound

This paper cites Rouge: a package for automatic evaluation of summaries,.

Engineering AI Judge Systems Rouge: a package for automatic evaluation of summaries,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.968535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.330973Z digest=sha256:4dbf063e66833d3e812b5faa4452acaab58acc14baa0bb1d25e311f6a423d62a

Observation 56a5924d-8edc-4c5f-8af5-dc72367b80be · outbound

This paper cites Best Practices and Lessons Learned on Synthetic Data.

Engineering AI Judge Systems Best Practices and Lessons Learned on Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.334631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.334631Z digest=sha256:d351016afb503db97bbf37afccb4b832ba8fd23d573c471c986c3d9b623947ef

Observation a57f28fd-7676-4364-9c5f-b975c931315e · outbound

This paper cites Calibrating LLM-Based Evaluator.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:89ee2d1530c0a5326e450974ed6095a08192855021f7dc740d364a044805e744

Observation ae9b3172-cb8b-4264-b00e-b1fc5795ef68 · outbound

This paper cites Generating training data with language models: towards zero-shot language understanding,.

Engineering AI Judge Systems Generating training data with language models: towards zero-shot language understanding,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.956636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.343429Z digest=sha256:e41268deefe554e125eae2c9620753bffde0ff3e619dfbf1bd15d4a6a9db3600

Observation eb717707-4edb-4c03-aeb5-23bf499ef0a4 · outbound

This paper cites Cider: robust consensus-based image description evaluation,.

Engineering AI Judge Systems Cider: robust consensus-based image description evaluation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.944650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.348047Z digest=sha256:35566d2eb133604536f3c94eb5d21e095a8f7ffa70b9d6988f353c4e652b8564

Observation 95a55474-bc6c-4ec1-81f6-73192c62e639 · outbound

This paper cites A framework for evaluating and improving requirements specifications based on the developers and testers perspective,.

Engineering AI Judge Systems A framework for evaluating and improving requirements specifications based on the developers and testers perspective,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.932435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.352122Z digest=sha256:d781af27d4be99537e8de1d05b222fc9606adb86391d23a845cbeaecc59bd662

Observation 698c9c48-3d7c-4663-9190-028b2013d533 · outbound

This paper cites Human-Centered Design Recommendations for LLM-as-a-Judge.

Engineering AI Judge Systems Human-Centered Design Recommendations for LLM-as-a-Judge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.356623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.356623Z digest=sha256:a3b778669195ebd13b525240f216a7ffee9271351b01f0ff742c880f5103c1b7

Observation d380be56-886c-43af-a80c-d04386398aaf · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Engineering AI Judge Systems Bleu: a method for automatic evaluation of machine translation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.920316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.360737Z digest=sha256:fd3fe4e5136bfcb678c3832255837e8552b9f1a6fc10ed075210c39d74f3f98c

Observation a5a90d46-9e76-4ecb-94ce-4bf898a63156 · outbound

This paper cites The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities.

Engineering AI Judge Systems The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.365781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.365781Z digest=sha256:3eef11b8d4273fea754663fd86c59bcc2b7d5b84b980323a7b7d0522e1f32c04

Observation 57920572-3461-49cf-b1b9-5a97faa05cee · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Engineering AI Judge Systems Verbosity Bias in Preference Labeling by Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.369902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.369902Z digest=sha256:64a3e48853f853f0f4a564aee42fbfb7f14b0ea51009336d95c1c93d1f63ef81

Observation 0f1b10a0-d706-48ae-a6fd-e0e483db6f0e · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

Engineering AI Judge Systems BLEURT: Learning Robust Metrics for Text Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.374841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.374841Z digest=sha256:8c7b13e46dd4a7846763b72892cb8d6c4f53f206e5bde979962615cbc7fe5318

Observation 1cd8717b-f6b1-4089-aec4-d5b796dbcae0 · outbound

This paper cites On automatic summarization of what and why information in source code changes,.

Engineering AI Judge Systems On automatic summarization of what and why information in source code changes,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.906996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.379788Z digest=sha256:bae30bc4ed68f38ebf620936b0b9c73219c4de537cf1c6331b76eb1fefc9c180

Observation ff1d24d0-ab0f-4ea3-92d4-57d3d1841f88 · outbound

This paper cites On the evaluation of neural code summarization,.

Engineering AI Judge Systems On the evaluation of neural code summarization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.894287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.383721Z digest=sha256:d990d363aa622cdb30e557ced3facba2fa713bfe9e2700b6124ba052a16d875a

Observation f1f2be3f-b8ef-4e6c-9a72-9c183b1bb0f9 · outbound

This paper cites RACE: Retrieval-augmented commit message generation,.

Engineering AI Judge Systems RACE: Retrieval-augmented commit message generation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.882388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.388673Z digest=sha256:8d588e4bb48a583dd60e83440db5203ee4cd96a6d6e76a16bff56cd59a9b6843

Observation 9de57533-72ab-447b-9c2e-9a16d1936911 · outbound

This paper cites On the evaluation of commit message generation models: an experimental study,.

Engineering AI Judge Systems On the evaluation of commit message generation models: an experimental study,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.870426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.392463Z digest=sha256:2d7e8f235171733d9be63d004dd44ee5a8bcb3637c5dff8cc618dfcb58d0055b

Observation 68bbccb3-869f-49bb-aa65-0f68b69ef797 · outbound

This paper cites Alpaca: a strong, replicable instruction-following model,.

Engineering AI Judge Systems Alpaca: a strong, replicable instruction-following model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.858724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.397190Z digest=sha256:d43bd5cfb2eb94539d5cd76cf378c3ae22cb017b6bd90ff9391918d54d8a5fbd

Observation 196eee81-2147-4992-9d96-4f60defcddda · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Engineering AI Judge Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.400788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.400788Z digest=sha256:97d225d32629d3b44a5f50e87f2c338a270e713c025aaa3c8c80a5c88224db66

Observation d64aa7a2-c591-4f8b-8884-5cb6100a9cc3 · outbound

This paper cites Synthetic data, real errors: how (not) to publish and use synthetic data,.

Engineering AI Judge Systems Synthetic data, real errors: how (not) to publish and use synthetic data,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.844266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.405510Z digest=sha256:32fcc88db764beeae0bf13cc8d7b32c210800600c8335159b3c51125f88d68ab

Observation a7d2c3be-6852-4eea-b3be-9e62162e1d76 · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Engineering AI Judge Systems Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.409126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.409126Z digest=sha256:0ab3ff1ef5f12e86e60741ad8b59bc6a9e2480b269d55df806de7b3744d8acc9

Observation 02d0b65c-f5f1-4087-b51f-22bad0f786a2 · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Engineering AI Judge Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.413886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.413886Z digest=sha256:06b54b060c1434d6d05a3f85c488ec8742226d49b8493162301bedaa1ab72dbb

Observation 5746f32f-31dc-40ec-b425-a3aa15be497f · outbound

This paper cites Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems.

Engineering AI Judge Systems Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.417892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.417892Z digest=sha256:51bc2a42d0fb250afe5b9dfde947a90c406bdf624e8137e4795a0ae04d364094

Observation 99eedc02-be11-4d1c-bb22-864eaef6e7e4 · outbound

This paper cites Fake it till you make it: face analysis in the wild using synthetic data alone,.

Engineering AI Judge Systems Fake it till you make it: face analysis in the wild using synthetic data alone,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.831987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.422862Z digest=sha256:38000b1388f15dab8108635295541eaf830b228cb659acd5081079a3cc11585e

Observation 219aebda-04d4-455d-b19a-6145c4cbea0a · outbound

This paper cites Commit message generation via chatgpt: how far are we?.

Engineering AI Judge Systems Commit message generation via chatgpt: how far are we?

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.820449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.427237Z digest=sha256:bece3ac6bfbfa4c21c2d2f995692fdd934fe20e4f16c41a7d5d10b8cca6a9200

Observation f890a40e-3532-4f3d-b0f0-0e99dbe9cb2c · outbound

This paper cites Automatic commit message generation: a critical review and directions for future work,.

Engineering AI Judge Systems Automatic commit message generation: a critical review and directions for future work,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.808693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.431034Z digest=sha256:deec30ede41061a2e0ca6e948a6f1811bbcb7734f4f191b4b9061b661e85b0b3

Observation b3e532f0-e1c9-45c0-8e81-6851a8f47262 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Engineering AI Judge Systems Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.795757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:59:18.434546Z digest=sha256:96da099e570379b4a85873038e20481fc2373b03c015087145aeb416d01037e9

Pith citing papers

Observation a1f08563-e1d9-450b-b1fb-942ab696aa28 · inbound

Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers cites this paper.

Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers Engineering AI Judge Systems

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:32:15.515399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T09:32:08.596259Z digest=sha256:b4f0a6dd70cdd7982e3160d6dafa002cf58a482910532a0cd4117ca359317969

Observation c3ec2d89-6bea-4e6b-8f3b-09f152b10195 · inbound

An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges cites this paper.

An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges Engineering AI Judge Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:18.383835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T10:32:47.756343Z digest=sha256:c1704c675e0212023d2406916ea54f4f168b55a8852c4d6f125e81c70c611725