Pith. sign in

Paper Citation Record · LEDGER

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems

As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.05638.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05638 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T04:31:30.304599Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c67e26e7-20ab-4af6-9bbc-0f930f16ca4c · outbound

This paper cites Tracking the moving target: A framework for continuous evaluation of LLM -based tools.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Tracking the moving target: A framework for continuous evaluation of LLM -based tools

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:5464a95dcc074c9d96043453ae4a9a9ecc924d2e1331e699d9d88ad8cae6049a

Observation 8122f513-1b5a-49fb-aa0d-1a8a6efde94b · outbound

This paper cites Sales research agent and sales research bench.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Sales research agent and sales research bench

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-11T04:37:48.195193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:22dfc0791e986e02c69ee8f600e11ed7c0409dceafe3fe0dc8a30d840d234e65

Observation dc950f7c-6f1a-4df7-b2f6-293941ecf39e · outbound

This paper cites A survey on evaluation of large language models.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems A survey on evaluation of large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:0a3fded5f14013189f29b56b2da7cf58e88404f4e62ad456f95f999cfd816c84

Observation 7eeb8257-39f6-4b3e-819c-4be7efdba038 · outbound

This paper cites CritiqueLLM : Towards an informative critique generation model for evaluation of large language model generation.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems CritiqueLLM : Towards an informative critique generation model for evaluation of large language model generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:ea34af9fa19e0b9644afdab3a82f7cfa44da71887c0edcd652a293ff38b47dd1

Observation b0949089-828b-4d80-a664-8c01a1f0aca5 · outbound

This paper cites DSPy : Compiling declarative language model calls into self-improving pipelines.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems DSPy : Compiling declarative language model calls into self-improving pipelines

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:84e9b46659157b1f1554a3f8f84bd9618b31a52a12ecacc77d1ac758cb08a523

Observation 72d2d7b1-28f4-4e7c-a06e-d6a9104b01ce · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Prometheus: Inducing fine-grained evaluation capability in language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:17f2aa3a3c2bb64397aefa670705e3771bb5e770abdf7afa74c407333e7677d0

Observation 6f4470e7-f9ca-4a75-b77a-ae60e340d52b · outbound

This paper cites PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:cb8fb32daa15dfa9fd116d9491cccb483c905265efa08602937a2224b837b008

Observation 40fce22c-1a0f-4be1-af4d-249beb098c64 · outbound

This paper cites Holistic evaluation of language models.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Holistic evaluation of language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:49d56eead262a335fdadd82bc369c3e3726611f3dbfa15f668ae8f899569567f

Observation 7d3553b1-3741-45b7-97d8-a1606df935bc · outbound

This paper cites G-Eval : NLG evaluation using GPT-4 with better human alignment.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems G-Eval : NLG evaluation using GPT-4 with better human alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:8ab72717f890a16e050c83221a36b22c065cf2ed0bd0adbb826d01712a559774

Observation b99cbb13-fc7c-4a5b-83cb-1a8d7b595122 · outbound

This paper cites FActScore : Fine-grained atomic evaluation of factual precision in long form text generation.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems FActScore : Fine-grained atomic evaluation of factual precision in long form text generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:801026fd669d8085781a2c30f6cf30c5bb777f1efde2a0e28818e6815218c040

Observation fb9b4ffe-3174-4067-8e82-d2a3d73095a1 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:73ee5d0f8cd2e53b3bfcfc22942eef52909a47561e3ca3217a226205b7640971

Observation 8565a891-a4ab-481a-bbee-d694978ca86e · outbound

This paper cites Continuous benchmark generation for evaluating enterprise-scale LLM agents.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Continuous benchmark generation for evaluating enterprise-scale LLM agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:ecf9f0516b9d08641d208028467747895feda8230a46752611d409406026c17a

Observation 50be605d-b448-4310-a15f-be876f45bb51 · outbound

This paper cites Quantifying language models' sensitivity to spurious features in prompt design.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Quantifying language models' sensitivity to spurious features in prompt design

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:6f3d719414650ded0018df593c179f901612b93413766b4820f4c28173a79e5b

Observation 17b1b9e6-75ea-411a-a84c-3bb1236c9232 · outbound

This paper cites Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:c65a257c45b865ddeacf30b4e5bc9dc42b2c195957a1b911957b62bb29a59992

Observation 2df74d01-d23a-47e7-bc98-2185f8bbdf9c · outbound

This paper cites DecodingTrust : A comprehensive assessment of trustworthiness in GPT models.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems DecodingTrust : A comprehensive assessment of trustworthiness in GPT models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:eaa11825543009fa370876275f8fc06d6c60890ec0b090bf73196b14b572b97d

Observation a7252636-13b0-4af4-bb2e-54afe157cff0 · outbound

This paper cites Enterprise Large Language Model Evaluation Benchmark.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Enterprise Large Language Model Evaluation Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-11T04:37:48.225034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:695bd2f7ffc4bc3c0359a8415a3afb76e52cb1b9d8ab51006e0e58a65a0093f2

Observation 355fac38-aa57-4548-9269-bc4f155488c4 · outbound

This paper cites Large Language Models are not Fair Evaluators.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Large Language Models are not Fair Evaluators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:a9c764635388ac0593c6751df4d758e1e24cbd1fcbc3b99c57ce20ee1fde0ca7

Observation 5271d88e-1c75-498e-8765-04b134aae86f · outbound

This paper cites Enterprise Benchmarks for Large Language Model Evaluation.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Enterprise Benchmarks for Large Language Model Evaluation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-11T04:37:48.237015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:bd7ea28c31ef1c8cc9fa1a83948556f6890ffe3ab01a065df27c036b841442c0

Observation f86fe046-35ca-462b-83a7-6a2b18d42b0c · outbound

This paper cites Judging LLM -as-a-judge with MT-Bench and Chatbot Arena.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Judging LLM -as-a-judge with MT-Bench and Chatbot Arena

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:006b41679610feb1f9ed741556d99cf5a1bd69ab78a1a169c7bdb80e757620cf

Observation b9fb2545-7687-4857-a08b-96caf623a160 · outbound

This paper cites Large language models are human-level prompt engineers.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Large language models are human-level prompt engineers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:24eb17d266368cacbaea3a688aefb30b93dc744a25a8527c2c29b04ea0425e9e

Pith citing papers

No inbound Pith citation observations are available.