Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 7 inbound Pith citation observations for arXiv:2507.21028.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21028 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:07:16.662207Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:10:30.862323Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.866103Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 716e44de-a45b-425e-81ec-33135f596d20 · outbound

This paper cites Children - Emily Thompson Demographic Info: A 4-year-old girl from Seattle, living with her parents, attending preschool and developing strong language skills.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Children - Emily Thompson Demographic Info: A 4-year-old girl from Seattle, living with her parents, attending preschool and developing strong language skills

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:18.756411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.474260Z digest=sha256:b9cca0c73ff582cddb2c7cefb06198adf86130129def7a5ba78fc134b4756596

Observation cafdf5d4-6789-4ec7-b9a2-d7c293089ccc · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.589727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.485997Z digest=sha256:16667863c204851f97f7ea6905fc5b23be78cb81e36a20862cd583caca560e82

Observation 1e236875-3995-4f26-8206-c7be88fe7143 · outbound

This paper cites [The Start of Assistant 1’s [evaluation content]] Clinician - Dr.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation [The Start of Assistant 1’s [evaluation content]] Clinician - Dr

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T13:07:18.491719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.496469Z digest=sha256:fc8f5a22a4dff7bb341e93333cdf3c023498f5b0fd852e50900343126a7957dc

Observation 3bc51568-58a8-41c4-90e6-d29a5b35d111 · outbound

This paper cites stakeholder name.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation stakeholder name

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:18.067418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.538905Z digest=sha256:132de94ad3ac4cef1fc72eb21d6e7adc7933011ac81f8743a7cd103533fc635b

Observation fedf7732-ddfc-4f6d-b1c8-f1916c3b6c93 · outbound

This paper cites RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.449854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.449854Z digest=sha256:f3f5c126a2c27fee96ec506ca54f9b1cd6039f10d9b493cedf60f832995b0061

Observation 3098da55-bbef-4d16-8d7d-bb40b5170521 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:07:16.461555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.461555Z digest=sha256:32ef5d39f0f340233032ee07430bcc162c1b7723604e5b77d160b57a40b53740

Observation dea287df-0d39-4a14-ae96-f922a9307eed · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.397249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.506924Z digest=sha256:291b0b95bdd74a8fa46421050a48cef3805386a796f13b53c74b1d037bb1cac0

Observation 9805c255-0191-43a6-9321-c19360f8d55c · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.262717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.517819Z digest=sha256:34a048545463295ec6eb88179ad5901f8cb7fc7c0e0a3eb758e1d5bdc298919e

Observation a7a8cda1-9b5b-4618-bef5-bcc3a7b28e95 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:18.169335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.528319Z digest=sha256:2710aa843c23dae84f3b3a71967e20573d2ca89c2544a2f3bc8461eb76f92e8e

Observation b9f1750c-aa7e-41d7-8bd0-e405bddc597f · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.936064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.550550Z digest=sha256:d0573b3596f316f02e49d2aa28bc3256986ef0740a73cedf85163f2680525ec6

Observation cf8ca2dd-abba-4100-a234-6543f5b72cd3 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.828232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.566947Z digest=sha256:bda36fa878cbabb42a6f00f58ba2df769e9c1f0640b41e45fa8655b1994bc7f3

Observation 8a89f0eb-e735-4355-b849-3ac8c3bd77ce · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.709490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.582300Z digest=sha256:e19998862a42cafb6c0c59ebe985c61dbcf2fb29db9f80e4c88e3c0a8d72a032

Observation 9987aec2-8cf4-4b6b-ab7d-9a6cedcf1744 · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.612913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.599811Z digest=sha256:aa8a2815e11cd70bda8dcbafb87e8dfa5c24eb3c3531bfca8756cfb3ba6f7697

Observation 85450635-3253-4409-b20f-ecfe71643ea6 · outbound

This paper cites Stakeholder Name.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Stakeholder Name

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.514777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.615269Z digest=sha256:1ca8016edcc2be47606f3f3e6dea8159ce94711aaa4cdadde289339218698f96

Observation ce1f668d-c750-4690-b07f-2f95274f4cab · outbound

This paper cites an unresolved cited work.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:17.419480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.629445Z digest=sha256:5e8b2063cf3ec4417c1326cfd16a3bfb97df9221cc694610ee882c006f4ad9ee

Observation 9c348cef-30ef-4152-9936-fc958c9ac100 · outbound

This paper cites NO MORE COMMENTS.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation NO MORE COMMENTS

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.291123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.643778Z digest=sha256:da23a664afddf932f6f47be322e40be751c211432f28c2b6555a725f1395e2e5

Observation 32637be4-9490-4266-95fa-00e89f0511fd · outbound

This paper cites Here are the initial evaluations from all stakeholders: {phase 1 evaluations } Your task is to evaluate these initial assessments based on your perspective and/or specialty.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Here are the initial evaluations from all stakeholders: {phase 1 evaluations } Your task is to evaluate these initial assessments based on your perspective and/or specialty

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.197620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.651451Z digest=sha256:2df74fa4d7c36d9bf80b0e86c7873b5b695e60f80f7863e40be885fce27740cd

Observation 680101d2-51c8-41b8-86bf-ceb04b3fcae2 · outbound

This paper cites NO MORE COMMENTS.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation NO MORE COMMENTS

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:17.034612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.662207Z digest=sha256:6a2a019a5e997b6e8e3272843079234558fe86a504acccb4c9c2255e63a7d4f6

Observation 4ae65409-9274-4001-8058-71cc0a4b299b · outbound

This paper cites Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:07:16.926691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:07:16.413317Z digest=sha256:e95d3797f32ec1f5950b06607a054346fe3be47ef80257884922bcc252b611ae

Observation cad370de-dd9e-4e7e-9ffd-0ff29d511c6d · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.404481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.404481Z digest=sha256:6f73e723d4acf7332ab1aa860f0150409ca6057ccad231660bb58c4c6fd50c23

Observation 31f1bc9f-2122-4687-894d-ac7021079360 · outbound

This paper cites LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.435775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.435775Z digest=sha256:86f54c4a5c05b722b35774ca11e49d1563f43b60ff9c3f551d189758e591269b

Observation eaa8a27d-a441-4481-bcb0-058c0b67b85c · outbound

This paper cites In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, New York, NY , USA.

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation In Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, New York, NY , USA

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:16.422913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:16.422913Z digest=sha256:136ef1e3d9206549ba0da8bd63cc6652e6f3e18aba4fa053cb02f6a1ffa02ace

Pith citing papers

Observation 9a198458-32f2-4503-8984-9026ed1cded6 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.685017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:e4847d51adb686217fc729f34db552062e23796b64544176e18bc2440cac2ebc

Observation d7e3ba71-1df1-4d01-bf4d-b237be20aaa5 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:10.316870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:2ac48cf4c4a186ec0e3ab8f5b90667a75af32a91748567a3a4bbd483f8f35af3

Observation cf1850b0-462a-4e6e-8579-da1067759e4c · inbound

Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework cites this paper.

Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:06:34.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T03:06:05.168589Z digest=sha256:d3ffce29b21faf5f27755b6e27bb1276bf271a056b0827596373578c8a3bb5dc

Observation e0d301f2-7ff7-4fe7-8bdc-b8f57a1eefc6 · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.750906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:366a8697dbe7875af4f7c42309d9d8958fa6c74531e67308f9612a2a6a69cb01

Observation cfb32351-f75a-4e78-8076-1e30c3cd3f09 · inbound

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments cites this paper.

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.867675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T01:25:51.788544Z digest=sha256:ec5aa25b80818e4f8f1040b1b35e9b1bf3db9736506c05ea6131530098cdedd6

Observation aa634f94-6298-4de8-8bb9-44b4a3b196de · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:36.496381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:36.496381Z digest=sha256:cfb059bf2f95e094c489f8ba2116d9b2e1a0ef3a56839fe43a99dc892471a136

Observation df871948-4df4-4ede-9831-4217e1e92d9b · inbound

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces cites this paper.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.862323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.862323Z digest=sha256:82ad3b7c54c20bf2b3c51a51bbbce76549ab830953bc8a4b69eac094faf1ac89