Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T04:31:30.304599Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.05638.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T04:31:30.304599Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c67e26e7-20ab-4af6-9bbc-0f930f16ca4c · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Tracking the moving target: A framework for continuous evaluation of LLM -based tools
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8122f513-1b5a-49fb-aa0d-1a8a6efde94b · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Sales research agent and sales research bench
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc950f7c-6f1a-4df7-b2f6-293941ecf39e · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems A survey on evaluation of large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eeb8257-39f6-4b3e-819c-4be7efdba038 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems CritiqueLLM : Towards an informative critique generation model for evaluation of large language model generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0949089-828b-4d80-a664-8c01a1f0aca5 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems DSPy : Compiling declarative language model calls into self-improving pipelines
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72d2d7b1-28f4-4e7c-a06e-d6a9104b01ce · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Prometheus: Inducing fine-grained evaluation capability in language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f4470e7-f9ca-4a75-b77a-ae60e340d52b · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fce22c-1a0f-4be1-af4d-249beb098c64 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Holistic evaluation of language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3553b1-3741-45b7-97d8-a1606df935bc · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems G-Eval : NLG evaluation using GPT-4 with better human alignment
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99cbb13-fc7c-4a5b-83cb-1a8d7b595122 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems FActScore : Fine-grained atomic evaluation of factual precision in long form text generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9b4ffe-3174-4067-8e82-d2a3d73095a1 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8565a891-a4ab-481a-bbee-d694978ca86e · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Continuous benchmark generation for evaluating enterprise-scale LLM agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50be605d-b448-4310-a15f-be876f45bb51 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Quantifying language models' sensitivity to spurious features in prompt design
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b1b9e6-75ea-411a-a84c-3bb1236c9232 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df74d01-d23a-47e7-bc98-2185f8bbdf9c · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems DecodingTrust : A comprehensive assessment of trustworthiness in GPT models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7252636-13b0-4af4-bb2e-54afe157cff0 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Enterprise Large Language Model Evaluation Benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 355fac38-aa57-4548-9269-bc4f155488c4 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Large Language Models are not Fair Evaluators
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5271d88e-1c75-498e-8765-04b134aae86f · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Enterprise Benchmarks for Large Language Model Evaluation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f86fe046-35ca-462b-83a7-6a2b18d42b0c · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Judging LLM -as-a-judge with MT-Bench and Chatbot Arena
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9fb2545-7687-4857-a08b-96caf623a160 · outbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems Large language models are human-level prompt engineers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.