Pith. sign in

Paper Citation Record · LEDGER

LLM Critics Help Catch LLM Bugs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2407.00215.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00215 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:55.260097Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bbd0236-5e0d-4e77-9466-a2c223a80c5e · inbound

Automated Capability Discovery via Foundation Model Self-Exploration cites this paper.

Automated Capability Discovery via Foundation Model Self-Exploration LLM Critics Help Catch LLM Bugs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:55.260097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:55.260097Z digest=sha256:28a2e510667dbd6551adfd26ffd6249c84edeb31ded478a1ff38432ef72bfee3

Observation ee956037-fb1a-4883-9db4-5f2edfa415a8 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models LLM Critics Help Catch LLM Bugs

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.064965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2843af35d51e4e9c55e1cbe995f1753cdb6fa70c18891d374a34bb0c2cba8a31

Observation 71262754-d894-4f4a-a2d8-3c63aa8ffb1a · inbound

Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development cites this paper.

Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development LLM Critics Help Catch LLM Bugs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:36.041322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:36.041322Z digest=sha256:3d67406da80e950a33d15bf10b76025238f603849fd8d03599c6439566b8ec49

Observation d664ad6a-ee01-489f-9aac-0f635ddb0cc3 · inbound

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents cites this paper.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.997683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.997683Z digest=sha256:131689a08bbdafb053b3e4b5793e33581c61eb70f7f4ab391f4151c15c043e40

Observation b36489cf-d06b-4897-b3c9-75ceb9958269 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review LLM Critics Help Catch LLM Bugs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:29.658971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:29.658971Z digest=sha256:9f80bacf74cc9290fc407f587cf4e417c9e52241dec4a84dd9a854660230e14d

Observation b97f700a-87c7-4fa0-8cb7-7840c538bd65 · inbound

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models cites this paper.

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models LLM Critics Help Catch LLM Bugs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:00.096836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:00.096836Z digest=sha256:97decc172ae21cf3affef01b060946ce5014d743a7d7559ef2aec8d2056d3c39

Observation 12932292-e91b-4517-9c01-14b43c65cf96 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations LLM Critics Help Catch LLM Bugs

Reference 261

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.148271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.148271Z digest=sha256:aebc5a959c090d84288b2b431d3af8ac52d52d40bb0cf2a9e1f863c0fc5d0a93

Observation 02bc1dfe-46d6-417a-bbef-fef24ef561b9 · inbound

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios cites this paper.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LLM Critics Help Catch LLM Bugs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.053326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.053326Z digest=sha256:f8b64972078d288bccd5d1cfe02c6a4a147eafa64201dc4c0d02c395ea5d378b

Observation 7fae0959-5729-4815-b117-045361ba19bc · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLM Critics Help Catch LLM Bugs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.764833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.764833Z digest=sha256:16d229557bff86e4d702dc421d8bc522171142d0394e931ab5e61ad5453ba56e

Observation 6868fb9f-f436-4f7d-959a-d7d92e1ba5c1 · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks LLM Critics Help Catch LLM Bugs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.445876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.445876Z digest=sha256:b58d5a3d07f903e042cc46f3acd6cbb80771ca462e27f71a2154c253928c9b66

Observation d82d28eb-19e1-4866-b18f-23d8360e1bd6 · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety LLM Critics Help Catch LLM Bugs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:19:44.746613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:1891ab94aa0039f082ec8700191aa17c608b2596038a5cd0efe3a0971981103a

Observation 395ff2a7-3a94-42ad-afc1-0e52bc8d9e71 · inbound

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback cites this paper.

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback LLM Critics Help Catch LLM Bugs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:32.141323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:47:32.141323Z digest=sha256:2b224d7ee8858732e9f8dbcee713fcc57ab77251114dfa9630fc28ba973a9e72

Observation 9adddc11-31b6-4e55-a087-ba9a64a6eb06 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning LLM Critics Help Catch LLM Bugs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.268233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:cc72458df955414ccdea81476126ef4677a4070a134ea390a8acfd643a345e39

Observation da4951f9-e021-4c94-b175-b9086c68171d · inbound

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts cites this paper.

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts LLM Critics Help Catch LLM Bugs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T05:46:52.666443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:46:52.666443Z digest=sha256:5916e3899db1d1e0d4d602b9b3957171c76ff4ec3640e6401d9ba626f5a1db61

Observation b0539d43-abd0-461c-8d1c-864fe2e38c9e · inbound

Are Today's LLMs Ready to Explain Well-Being Concepts? cites this paper.

Are Today's LLMs Ready to Explain Well-Being Concepts? LLM Critics Help Catch LLM Bugs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T01:02:51.111485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:02:51.111485Z digest=sha256:e89bc6feb577f8ce05f742c817fd52928b80ddc10c809e42f1ad9b3074d62259

Observation 48faffc0-1935-46f9-a8d2-a314fba9378f · inbound

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs cites this paper.

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs LLM Critics Help Catch LLM Bugs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:46.553605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:09:46.553605Z digest=sha256:fe16761731a6ba3d8c25b4e818c2a92693c75dd640b402b365422ed7c8ec8b34

Observation 8a10b93e-dbd6-43be-9e52-76ceea82646f · inbound

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models cites this paper.

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models LLM Critics Help Catch LLM Bugs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:38:06.861926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:38:06.861926Z digest=sha256:09d7943dceac01425c19c0124d19e01da17b045a41f3c2e3f6853649fe629af3

Observation 007db2f9-132e-4e26-8ae4-18318fdfbb5d · inbound

Dream-Coder 7B: An Open Diffusion Language Model for Code cites this paper.

Dream-Coder 7B: An Open Diffusion Language Model for Code LLM Critics Help Catch LLM Bugs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T12:56:35.136328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:56:35.136328Z digest=sha256:9a8e6498915ae79874577e92f30f03ec08e70fa476453a7590de3e53fa76f678

Observation 983d0e01-9175-409f-93ce-b594b6db0d16 · inbound

Human-AI Complementarity: A Goal for Amplified Oversight cites this paper.

Human-AI Complementarity: A Goal for Amplified Oversight LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:22.511988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:22.511988Z digest=sha256:04262bd0c8038d21d54d8b039f7f40e83f23924adcfd8cec1502571e43940083

Observation b1730ac7-4802-48f8-9c5f-ae186195875e · inbound

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning cites this paper.

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning LLM Critics Help Catch LLM Bugs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:03:04.279092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T16:01:48.789986Z digest=sha256:dcada1cc32ae916fda68fdbf10d90cb2f051aad53df4456bc80e9c2fe304efe1

Observation 28fe0286-729e-4c34-97ad-3e29aef28991 · inbound

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories cites this paper.

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories LLM Critics Help Catch LLM Bugs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:36.059248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:29:37.642356Z digest=sha256:bafeea700d785bab8bfe0f12124c79d94a89635cd1349de03bccdac72bf7c210

Observation a5806370-9914-4158-9a1a-0b5fc8847c67 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight LLM Critics Help Catch LLM Bugs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.557005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e466e452a35a66c8eb08597b000414ec0171956003e1e18b75bb4fc5787b2aa1

Observation 954bd355-01ed-4240-a08c-369ca313ce8e · inbound

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks cites this paper.

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks LLM Critics Help Catch LLM Bugs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T00:24:28.369167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:30:55.419225Z digest=sha256:24deb205c5dc06ad57e691471a25ec3fabfc4c8184450d1ec89c1d4c0b9b2512

Observation 0435cd55-1b72-478f-b536-62a92a510796 · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction LLM Critics Help Catch LLM Bugs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:10.325818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T14:07:44.717260Z digest=sha256:a7c1d356c150bde67bb060784a252dc357e10f6205ec7334ee5bf2188562d956

Observation 647e6ab5-a8ed-487c-80aa-51a9564a395b · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction LLM Critics Help Catch LLM Bugs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:26.693683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:51:31.357544Z digest=sha256:9460eb7b8d0841513abdb74898cbd192adf9e5130efd696aa6ff9941488ebef6

Observation f47c57aa-73df-4737-95fe-b9752c8cb113 · inbound

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight cites this paper.

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.528520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:51:17.434889Z digest=sha256:30fec76e4c404d2708bc6e1122532d54f6b51a6721427900c14bcb55c4818366

Observation 21af213a-b98c-4ed4-820b-f4daf76d597f · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment LLM Critics Help Catch LLM Bugs

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:58:28.867558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T01:58:19.295247Z digest=sha256:23c5c2534b026711e00702b4c7f4050c3d1e3dfa96027b8d71939360ac64be93

Observation fd0fc45f-6ddd-4935-a697-a8f60dc628b8 · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment LLM Critics Help Catch LLM Bugs

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:47:40.465434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T16:45:16.963802Z digest=sha256:e73772d111de9c0a6fb4d8d65f83fd2ae9632f1644d05ca95f7768dd622c5c87

Observation 580e838f-3c39-49ec-ae90-046e36b15931 · inbound

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study cites this paper.

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study LLM Critics Help Catch LLM Bugs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:10:22.403624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:07:12.832683Z digest=sha256:8144c5f84cb1604fe0ec09a9700ace751251492ad1ea7919b56238d215b94dc6

Observation e8979a3c-728b-483d-bbb9-c65035c546bb · inbound

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight cites this paper.

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight LLM Critics Help Catch LLM Bugs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:10.983153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:55:58.645062Z digest=sha256:fd2925a3910ca994b4540c4bc75e4f774b4908b5c324e8c69185c2682dda4106

Observation 75ea0ca6-6ce5-4326-9f02-775fcb3bd834 · inbound

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control cites this paper.

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control LLM Critics Help Catch LLM Bugs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:50:39.714101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:50:39.714101Z digest=sha256:7a07c2815a23f15c9b5821ee4b75856ca1aa4530c84b5d2877e78d6dfd68604f

Observation ee3d1936-9e66-4f5c-933f-b42f661aa8ea · inbound

Fantastic Adaptive Taxonomies and How to Use Them cites this paper.

Fantastic Adaptive Taxonomies and How to Use Them LLM Critics Help Catch LLM Bugs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:14:42.470394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:14:42.470394Z digest=sha256:cd73b75800a865d366f4dc988b20346c64927a1cfaa58cedd44a6eca46c2d581

Observation 5a5900ed-38ad-4d3e-b70c-10be22d5ff0d · inbound

Code Monitor Red Teaming for Public-Test-Passing Code cites this paper.

Code Monitor Red Teaming for Public-Test-Passing Code LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:25.169891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:12:25.169891Z digest=sha256:f652b622ee3ffa378a44ca11c4afe492e5a6b4de7410f28f998e22118d86ca5c

Observation b5b8ce2a-122f-47d0-b665-1c83909c0789 · inbound

A dataset of rated conceptual arguments cites this paper.

A dataset of rated conceptual arguments LLM Critics Help Catch LLM Bugs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T01:28:21.684871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:28:21.684871Z digest=sha256:a2927045c2b701b9a17a3934060525b803cd3748eb3008ea86c49b20cd766984

Observation c37075e3-5fbb-44c3-91c3-1f907f29e722 · inbound

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation cites this paper.

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T07:20:10.899680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:20:10.899680Z digest=sha256:d9359feb5c0a1ed3246be50689412e518b54d0c0c91081787c6e4dc9cfc5b941

Observation 088a8d31-2fec-43ca-9875-79c8f24ad036 · inbound

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation cites this paper.

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation LLM Critics Help Catch LLM Bugs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T01:25:30.883047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:25:30.883047Z digest=sha256:028b7bc275620f0b9d15273895a5df7b624fc70243e3b0f28ba72c9097617ead

Observation 12947eb9-15c3-41cd-ac6d-73160fddd018 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets LLM Critics Help Catch LLM Bugs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.039128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.039128Z digest=sha256:463262f53db3614a949edc0cc0e0993a0c9ee597dc4095aedb4c24fe044dbd03

Observation 1b97cf54-a4ec-4b8a-a023-7d41cfc8b182 · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? LLM Critics Help Catch LLM Bugs

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:07.111956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:07.111956Z digest=sha256:791e9010e454b5ae52148d5ce7fca0ac9b1acf7f266593faa8813ad91856b164