Pith. sign in

Paper Citation Record · LEDGER

LLM-as-a-Verifier: A General-Purpose Verification Framework

As of 13 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 5 inbound Pith citation observations for arXiv:2607.05391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05391 v2

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T07:02:51.850836Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:59:12.973916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:12:15.992700Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f55c91b1-5d47-4fd1-b038-e60d40f7924f · outbound

This paper cites Scaling Laws for Neural Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Scaling Laws for Neural Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:674e3fbae0ffd221f176210c8706890c7c4876bc90226dcf7ba3dd636a1a4d9f

Observation 8ffa3366-e39a-4f28-a0f8-52935f8c285c · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

LLM-as-a-Verifier: A General-Purpose Verification Framework Scaling Laws for Reward Model Overoptimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:9a6fd782801efd6a0403ccdacf63291e8ad0e5c33aa6a90e293fa964d2376214

Observation 71e92306-613b-4935-a0ad-b753355627b2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

LLM-as-a-Verifier: A General-Purpose Verification Framework Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:6d3c94dd970a8055a4a5b88e5a1c92378cdcf7530ff19caa1f85abfd79eba97f

Observation b4c59a75-bbb2-4f71-a591-975a8e9f859c · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

LLM-as-a-Verifier: A General-Purpose Verification Framework Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:c7c0aadbb9205de703a0b82da41da3df3fcfa76717b6b6c1e8d44c6c34358f19

Observation ab9f7920-47d7-416e-8226-cf0ff9ec9e1d · outbound

This paper cites URLhttps://arxiv.org/abs/2603.04304.

LLM-as-a-Verifier: A General-Purpose Verification Framework URLhttps://arxiv.org/abs/2603.04304

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:5b6a254f61db09d322a617cba00c07477503dbc7a3ba088f429d173fea9c752b

Observation d1f15dad-7497-4caf-aebd-1852f3a138d3 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

LLM-as-a-Verifier: A General-Purpose Verification Framework Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3e774e7cc6f5e3ee54f21cb79ebd63fede1c78e4c0cf4d2647be575de621042f

Observation 128752d3-1f9c-4f5f-b981-a92d58aa5802 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLM-as-a-Verifier: A General-Purpose Verification Framework Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:0f679c9a97b4c78ab0a90e5f210a5199d88118a1f99c16ea33656992e6f9b4d8

Observation f1bb6b44-408f-49d7-a544-9c8c132a57dd · outbound

This paper cites Let's Verify Step by Step.

LLM-as-a-Verifier: A General-Purpose Verification Framework Let's Verify Step by Step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:aef714ad657fc7602807753004aef5b745cbf0e760eb7e8ad1a6d713bab7a84d

Observation 44dc11b6-73c7-4735-befd-005e2d059244 · outbound

This paper cites Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons.

LLM-as-a-Verifier: A General-Purpose Verification Framework Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:533ab670dd49c32db4b8b6637012da3d77bb765f8a6f5e4d365b3e19d002e871

Observation 3c6e1b83-dc0c-465b-8bd1-21eb50117013 · outbound

This paper cites Topreward: Token probabilities as hidden zero-shot rewards for robotics.arXiv preprint arXiv:2602.19313, 2026.

LLM-as-a-Verifier: A General-Purpose Verification Framework Topreward: Token probabilities as hidden zero-shot rewards for robotics.arXiv preprint arXiv:2602.19313, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:85d6ac9bc6396379e46c09a55ce02a50cb04b8456f089ce889c2ba442c955048

Observation fb8a223f-ef8c-4707-8c94-8c170bad9bf6 · outbound

This paper cites RoboReward: General-purpose vision-language reward models for robotics, 2026.

LLM-as-a-Verifier: A General-Purpose Verification Framework RoboReward: General-purpose vision-language reward models for robotics, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:f2eb71799194d05ae0007633f7fb6c14f61d6036f516c466af93fc9a3182e28d

Observation f5cb24fe-f1f6-496d-aeaa-a7c04d53b76d · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:de27cd36c93cbdbfbf9b00f741321d871f5fbebd90e980d66d9811e3175c2f10

Observation be5595fb-da5a-44bc-9937-ca43e97cf601 · outbound

This paper cites Steering Your Diffusion Policy with Latent Space Reinforcement Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:624f34b95e6381f7116e97084744f350f2150963bd5f2a49ded497a1121f3f05

Observation bcb07cd1-343b-40cb-9401-c4deeb814815 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:cb365f61c19ac7a2f7abd3d8621d750d67680abbe93effb72702ca40b093c80a

Observation 30f318c2-7c22-4f72-ae31-9b83b5d52cee · outbound

This paper cites Learning to summarize from human feedback.

LLM-as-a-Verifier: A General-Purpose Verification Framework Learning to summarize from human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3081728588ef1a6a81dc4b1435e220dfc2468a2d0891cc9e303e1b7e89979d89

Observation 4459bd1a-387e-4053-b86b-1f9915963925 · outbound

This paper cites Gemini 2.5 Flash.

LLM-as-a-Verifier: A General-Purpose Verification Framework Gemini 2.5 Flash

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:263f5919fc35cbfcd8cd6579e25c8ca284570e1382c3709847d663bdb57bb6f0

Observation 7d282e79-310d-4e78-88b5-5a15e58ec1c4 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

LLM-as-a-Verifier: A General-Purpose Verification Framework Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:e05525178cf31dcf24f080e468360103f244e1aa531774032ee1767873035333

Observation 7b4d97d1-5cad-4b99-b130-49daa6ab34e6 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

LLM-as-a-Verifier: A General-Purpose Verification Framework Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3fc1b2c42622e97684a6d9057fa5ba7935001339fd078494ebae19307791d6bb

Observation 7aa9590a-9cb4-4d3f-8a77-4db98162a74d · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

LLM-as-a-Verifier: A General-Purpose Verification Framework SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:52764683269c8ce71d53daa01167cfa2e627c3cdfa6d467e5141fa21e1e78268

Observation c3be7f12-5a67-4816-b980-af59e74381f0 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Verifier: A General-Purpose Verification Framework Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:9fdf33e776498fe833bddf338f7003ef09d00c5e6dc568933a657a047c2902c4

Observation 3aca4ddd-0a53-442c-b0f2-feea30e63f2c · outbound

This paper cites URL https://www.tbench.ai/leaderboard/terminal-bench/2.0/capy-build/unknown/ gpt-5.5%40openai.

LLM-as-a-Verifier: A General-Purpose Verification Framework URL https://www.tbench.ai/leaderboard/terminal-bench/2.0/capy-build/unknown/ gpt-5.5%40openai

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:047e9346409ebcd26f5796c8a6055c56bcb5063a259270850b2fa8a0ebf6f538

Observation e5e5f8e8-0a0b-43e0-a5ed-1629c0afb394 · outbound

This paper cites URL https://github.com/harbor-framework/terminal-bench/tree/main/ terminal_bench/agents/terminus_2.

LLM-as-a-Verifier: A General-Purpose Verification Framework URL https://github.com/harbor-framework/terminal-bench/tree/main/ terminal_bench/agents/terminus_2

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:fef06ccc01efb0ea0fb714e6d81c82645d60b6d1f37c70affe7081e3169dc517

Observation 27534865-1d49-4b99-b166-9facc21c7358 · outbound

This paper cites Vision Language Models are In-Context Value Learners.

LLM-as-a-Verifier: A General-Purpose Verification Framework Vision Language Models are In-Context Value Learners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:7dd15c7ef1a9aad11edb87060317bf14aa95feec3bd844cbb0d1fc542779242a

Observation b819d4be-6c06-42a2-8123-63425233c897 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

LLM-as-a-Verifier: A General-Purpose Verification Framework $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:632a72b9fe42ec2fd6c3673327f7a44bee6c7733b20b7a28b154c1cc721773a7

Observation bdcda714-7055-4a2e-a2a3-ecb6130f355c · outbound

This paper cites Qwen3 Technical Report.

LLM-as-a-Verifier: A General-Purpose Verification Framework Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:16f42600a8fec1757bf6dd6f4540d3d2bbbf981da18368b01b7b60591aac43d3

Observation b464b566-60bc-46e8-9b48-3048dbefe0f1 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:a9ede8d8f63e00a05d4cecf4146ef2bd00ae753cf844f4ad4a5c69eedaeccb32

Observation 93a2a312-e723-4b22-ac3a-a09547315ddd · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

LLM-as-a-Verifier: A General-Purpose Verification Framework Large Language Models are Zero-Shot Reasoners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:53ae32896ca62a65866d2d6b0d9bc347b9c6db6a3c2f7c353d17bc086f48cb79

Observation 1261d2b8-67f0-4493-baf7-e3dfde008ec1 · outbound

This paper cites Le, and Ed H.

LLM-as-a-Verifier: A General-Purpose Verification Framework Le, and Ed H

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:5b9a8ecd61d85790560c848a1bdfa820250eb3153427e5a230b6734a25a39766

Observation a63de5e2-2854-406f-bc2e-21e7271c45b4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:6437a8c987db8f8e0f6e8206238c8f11b13553ad20da7e3c3569e78ad55cf6e9

Observation 6ebb32fa-cd53-42ee-988a-bd4621eb7695 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

LLM-as-a-Verifier: A General-Purpose Verification Framework Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:fd280516f1515306305c2f963e426bf6b45cf15f24f5f07ccd263bef067296f2

Observation 5a6667e2-6171-4962-bdc5-b66d87e6ac1a · outbound

This paper cites Graph of Thoughts: Solving Elaborate Problems with Large Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Graph of Thoughts: Solving Elaborate Problems with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:242ea28733ae0439d0291741557825adb71bb52f11abc19b20fd87b267820c94

Observation 4851ab17-ff44-420c-ab3f-6cc4fd0aebca · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

LLM-as-a-Verifier: A General-Purpose Verification Framework ReAct: Synergizing reasoning and acting in language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:eeaac684fc3269dbd3e3927034299a2d8f72386891112337b57f26f7fefc7efc

Observation 89204035-561d-4e1a-9817-6978bcb41109 · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:c6fa3c9109a11ca3532de5e6b6280bd1d4f08ed70133788cc5a8089e73eec07a

Observation 5c6911e2-e729-4c75-aaf2-6cfcd0a7e6b2 · outbound

This paper cites Reasoning with language model is planning with world model.

LLM-as-a-Verifier: A General-Purpose Verification Framework Reasoning with language model is planning with world model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:73325a42a25d8f92a10cb65a31c31dc09cb9f4fbe2a736181edae520e2d3e4c7

Observation 23224799-541a-4ece-ad26-a35ac5686095 · outbound

This paper cites Large Language Models are Better Reasoners with Self-Verification.

LLM-as-a-Verifier: A General-Purpose Verification Framework Large Language Models are Better Reasoners with Self-Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:f840f39e0824fbc0128d5f5f49217894c20387dba214b006bdc68577019a9764

Observation 0c1aaa43-7926-405a-8b3c-4c7c23411194 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

LLM-as-a-Verifier: A General-Purpose Verification Framework Self-Refine: Iterative Refinement with Self-Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:d30dea43f9feea0e18c8aeeb8a9faae97600753946185e7055ce8c12bd2cb70a

Observation a511c53d-c90d-4b8d-ae4c-d0b5f6237d74 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:b1c9b4f95a5ff6f3be2d2907c91392cd0c6442554b12a1ed41ce4b152803edd9

Observation ae0aea5a-7598-46a9-97bc-e0e3860cc188 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

LLM-as-a-Verifier: A General-Purpose Verification Framework CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:eec32f9de0f67841b692431dbdd47b8ea9c2ff25e1083300b33090e901078ac7

Observation 5b7d6cad-5440-45f6-a52e-74e4dc9203df · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:ac8bfe721529083fc269e96a5e99d3c60b8acfdd093db0bbfa17c61e25a075fd

Observation c8b1c71e-98e0-4146-8c88-91dd8dbc4d8b · outbound

This paper cites an unresolved cited work.

LLM-as-a-Verifier: A General-Purpose Verification Framework Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:a5543954b252d2158b894e4e2b9ed39274b9e1f36415a7e222e4605c63dc91e9

Observation 5fed5ba3-8612-420e-9ab8-56de02cc4b38 · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

LLM-as-a-Verifier: A General-Purpose Verification Framework AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:f0befae5ac40a4d45602c6dd73054f6f1ab8d06fb14f1bda6bd26b4a92e0aeaa

Observation eac4cd1c-082f-4568-a1bf-2ecbec9cc440 · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

LLM-as-a-Verifier: A General-Purpose Verification Framework Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:23133f7c5e70066f010e1e90d0d93dd2f2804eea80b0531f1677ff1908d4f2ba

Observation abcc0918-3fa5-4e1c-8e08-62b5f903d970 · outbound

This paper cites Le, Christopher Ré, and Azalia Mirhoseini.

LLM-as-a-Verifier: A General-Purpose Verification Framework Le, Christopher Ré, and Azalia Mirhoseini

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3b129dca4e1059fc171257bfefb0ada0a7cd990b7b3ee48882499b9a62b22165

Observation 627b1007-72fa-4503-adde-915413945a20 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

LLM-as-a-Verifier: A General-Purpose Verification Framework Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:ca839d85459c1f72d847aa836e909ec5f119c0613b15b275de75919d83392e9a

Observation 0e2fcb5d-9bcb-4493-b8a7-56cd062cea4d · outbound

This paper cites Inference-aware fine-tuning for best-of-N sampling in large language models, 2024.

LLM-as-a-Verifier: A General-Purpose Verification Framework Inference-aware fine-tuning for best-of-N sampling in large language models, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:1b263c250d410ce4a3a3766c2d21fa01604a1360e85dd6703b018f15bf52111d

Observation b552de93-f21d-4aa0-85ec-2410ae7c14a1 · outbound

This paper cites GPTScore: Evaluate as You Desire.

LLM-as-a-Verifier: A General-Purpose Verification Framework GPTScore: Evaluate as You Desire

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3d4aba224de5d4121e27627f5c9bf5ebc8916d811241bb887d3313d8222c03e0

Observation 08f0cff5-7b56-4868-9730-75153130b0c3 · outbound

This paper cites G-eval: NLG evaluation using GPT-4 with better human alignment.

LLM-as-a-Verifier: A General-Purpose Verification Framework G-eval: NLG evaluation using GPT-4 with better human alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:9a22037cdf140c81487fe477f8ac8e1ad05954185bc24cea31019b14daada1bf

Observation 90437e6d-9db6-41f6-bc6e-72552aa43ca2 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.153.

LLM-as-a-Verifier: A General-Purpose Verification Framework doi: 10.18653/v1/2023.emnlp-main.153

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:613d11857236b80c232230869a5158c4f466222307fff835958f33b72a1499cd

Observation cb498ef1-45d2-4eba-939d-6b85ae9e122c · outbound

This paper cites Xing, Hao Zhang, Joseph E.

LLM-as-a-Verifier: A General-Purpose Verification Framework Xing, Hao Zhang, Joseph E

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:94879d05da02d35535ad754a92f8b5f9b3d333a250c45357c7f204c1eba7c772

Observation 6da08d49-f0c8-4d45-9b3b-e5799ab1a02d · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

LLM-as-a-Verifier: A General-Purpose Verification Framework Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:05ae7579597d664cfcac3ae0845f5610be1efee55d3ff95d034c6b9881958f68

Observation 2114fef4-2e7f-4078-bc10-19361bce13f5 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

LLM-as-a-Verifier: A General-Purpose Verification Framework From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:735a6b4e7e6073de39ba4469252e7a08ff728d48641002c63888369da9c8afc0

Observation 85b10f22-d2c3-4088-9cb6-b95c00ee20b5 · outbound

This paper cites FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets.

LLM-as-a-Verifier: A General-Purpose Verification Framework FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:acc7308ec7f9ae8ceffbc152d66268ab999f5d68741223f45de4d244ed4c4be1

Observation bb47685b-c3ea-4688-a947-2625490feec5 · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Prometheus: Inducing fine-grained evaluation capability in language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:7b126edb8db91482a643fd6fc4f9d367c25d0384d4c21cb51f18e4b46e040d30

Observation 65b118f9-6125-41c8-bc84-f3a60d53aa10 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Verifier: A General-Purpose Verification Framework Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:accf3a1afefdf6802903aa81cc23748bda8eb6b2c6108a4fb994cbefb2cbd60e

Observation 01659a7a-dfc2-40c8-a802-f63e4abeee92 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models, 2024.

LLM-as-a-Verifier: A General-Purpose Verification Framework Prometheus 2: An open source language model specialized in evaluating other language models, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:d2faa57d5dff1ddd091b7d31f1609774b3483b0d734aed5641308a31a7fdba71

Observation a8142372-901f-49c6-b35e-f893e46b74cc · outbound

This paper cites Generative judge for evaluating alignment.

LLM-as-a-Verifier: A General-Purpose Verification Framework Generative judge for evaluating alignment

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:24703e47574c9749678be99c8ccefa3cde24a9fe0e259815d324bb4ccd92205c

Observation 800f7d39-0df5-41c8-8f12-46c5e9f5be36 · outbound

This paper cites Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability.

LLM-as-a-Verifier: A General-Purpose Verification Framework Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:705a6c50e4971f2d3a335ccb50bd6c08ccc0fd89e60d58971e79c5420742e6b9

Observation 77cb4c56-63d1-42b1-bc9a-60a61b1a228d · outbound

This paper cites HD-eval: Aligning large language model evaluators through hierarchical criteria decomposition.

LLM-as-a-Verifier: A General-Purpose Verification Framework HD-eval: Aligning large language model evaluators through hierarchical criteria decomposition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:597d9696fad4465b78f15fc80556973b7eee763813d5631da8411e9ab9492eda

Observation d4c50f9e-03c3-43fa-bebf-da7e2fa5774f · outbound

This paper cites PandaLM:An automaticevaluationbenchmarkforLLMinstructiontuningoptimization.

LLM-as-a-Verifier: A General-Purpose Verification Framework PandaLM:An automaticevaluationbenchmarkforLLMinstructiontuningoptimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:b8d1113aa489e0b72c01b0a96ae354434c45c49f9a16511a799f7868e9a25884

Observation a299f0a0-01d0-4b89-bc8a-1a76daf7d99e · outbound

This paper cites JudgeLM: Fine-tuned large language models are scalable judges.

LLM-as-a-Verifier: A General-Purpose Verification Framework JudgeLM: Fine-tuned large language models are scalable judges

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:29cecd5f217bd0378ef3ae4db9605489e1ac31a24b642814312afe50cc7b67c9

Observation bf80e6ca-8834-4c90-9157-c6afb65304d1 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

LLM-as-a-Verifier: A General-Purpose Verification Framework ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:c2571a4ff7f1fa743e3aa2eb0cadf58c019f27024c88ac840deb8ed543cbabe1

Observation a5f1939b-7dbc-4939-b18d-c5fc4bb213c7 · outbound

This paper cites Shrinking the generation-verification gap with weak verifiers.arXiv preprint arXiv:2506.18203, 2025.

LLM-as-a-Verifier: A General-Purpose Verification Framework Shrinking the generation-verification gap with weak verifiers.arXiv preprint arXiv:2506.18203, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:e8176506f95b473725aaa0434ef65b856b2df82a0dc4ceddb64cbe6cf3732d9d

Observation 460285ab-49af-43f7-b100-9b5bc651eb03 · outbound

This paper cites Large language models are not fair evaluators.

LLM-as-a-Verifier: A General-Purpose Verification Framework Large language models are not fair evaluators

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:96f27e501a6a31ae8cf21b0865ab741517ca1477be7c9bad8849d9236e8b02c4

Observation 481c9dfb-0f74-4efd-961b-865f2bfc1593 · outbound

This paper cites Evaluating large language models at evaluating instruction following.

LLM-as-a-Verifier: A General-Purpose Verification Framework Evaluating large language models at evaluating instruction following

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:1ea2a6f332b97ef37c5a0117191878a472f812833b2b04b7a4d3d2517e7b2957

Observation c9b73bbf-03de-4ba3-b178-d538abc065d4 · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:5e0c18cb20a42c305260aaec065b3fc261a867086ed9f3d6cdbe1563291bd3b6

Observation 9d903bc7-1671-45df-82a5-b0383e87779d · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

LLM-as-a-Verifier: A General-Purpose Verification Framework LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:b262e8510324aa5735d832a463880b9023a918ea8e5073823417a812758a6822

Observation c1c3fb45-03e1-47ed-b3fb-afb656ed1620 · outbound

This paper cites JudgeBench: A benchmark for evaluating LLM-based judges.

LLM-as-a-Verifier: A General-Purpose Verification Framework JudgeBench: A benchmark for evaluating LLM-based judges

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:b954e19982766fadb27c2a0aaba06caa9d7ab2416d28386356104a5b901bea5d

Observation bad1870f-e62a-4db5-8c05-5f4d3eaa0e89 · outbound

This paper cites An empirical study of LLM-as-a-judge for LLM evaluation: Fine-tuned judge model is not a general substitute for GPT-4.

LLM-as-a-Verifier: A General-Purpose Verification Framework An empirical study of LLM-as-a-judge for LLM evaluation: Fine-tuned judge model is not a general substitute for GPT-4

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:1fb71db8c5dfe322f5312e4d97d27c6f3d95ab6faaaa142eaf72b201f283758c

Observation b7a9f655-9f30-4eb0-99d8-c68ab6d5b659 · outbound

This paper cites CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks.

LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:0c143c715d7bb6e953be97e0a3d8273e8e7af15300b3a63b9cc54661b31c6670

Observation 7fe4ef8e-7ef3-4cbb-90fe-4457cebebf50 · outbound

This paper cites MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark.

LLM-as-a-Verifier: A General-Purpose Verification Framework MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:a44601b940197a267d0be537bd2e40d6233666c2c9ef6c3a5f32e152ecfff962

Observation e071b0f0-db76-401f-adcb-695c60fd0a31 · outbound

This paper cites Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation.

LLM-as-a-Verifier: A General-Purpose Verification Framework Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:586c9f1333d973d1083f2ef58e0e1bc99c3e89a027da56bf83673a8bbc97bec3

Observation ad819da2-2bf2-4957-82d1-61e88aafbb35 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:916faeb397595791a6ba1c2d3767034087affad5be3d303e5bf044d8373eaa73

Observation 8c93d45e-0469-49ea-ba6d-17f975c76af1 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

LLM-as-a-Verifier: A General-Purpose Verification Framework Solving math word problems with process- and outcome-based feedback

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:768c30d993e9ca7f2ac96f255ddfbb31e17f75e485338aa59f2ef9a060fcf40c

Observation 2a02f34b-4e62-4d4a-9c0c-f39f55067a6f · outbound

This paper cites VIP: Towards universal visual reward and representation via value-implicit pre-training,.

LLM-as-a-Verifier: A General-Purpose Verification Framework VIP: Towards universal visual reward and representation via value-implicit pre-training,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:bb085fa84801939250b15cfb62db5f0a461cc2a46890b10e2e6a98aed7bcc4b0

Observation 1310acd4-f87e-4717-a0dd-aec57ce9dfdb · outbound

This paper cites VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training.

LLM-as-a-Verifier: A General-Purpose Verification Framework VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:b92e2c2fbe371ffe773ee78945f56b95417c5b39b6ae75977b595a3cb7fe6534

Observation a68434fb-7d6f-49c6-b07d-c8c18e3f2a6a · outbound

This paper cites LIV: Language-Image Representations and Rewards for Robotic Control.

LLM-as-a-Verifier: A General-Purpose Verification Framework LIV: Language-Image Representations and Rewards for Robotic Control

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:bdf753f81d528f10cd13418e966abe7d1500398cde06ad13d1c05de233cea198

Observation 81243133-fb41-45de-8cc8-fca2b085034b · outbound

This paper cites Sontakke, Jesse Zhang, Sébastien M.

LLM-as-a-Verifier: A General-Purpose Verification Framework Sontakke, Jesse Zhang, Sébastien M

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:ced70b8fe8d42831badb74ab5504053ac42f716556b7af1ba32cacf4f30de799

Observation afa61f81-4c7d-4877-8f5b-10580d8c0c8f · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

LLM-as-a-Verifier: A General-Purpose Verification Framework RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:5662c980344b1c171fc6c5efd8f495729e1cb8099339ca2a9b57323149ed5cc1

Observation d27b0c92-b711-49ec-a398-c26a00ba82a1 · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:ac646b925449fb031bcdfb38d5aba19bc91debef7e0968ed20ae070a94020f92

Observation 2159aea4-c9d6-488c-bf68-f1ac3e07fd84 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

LLM-as-a-Verifier: A General-Purpose Verification Framework RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:0a389cf15a991543d343ee9113ea562f32e97c73d601cd75e1e8f9badc6886b1

Observation 0e5994b9-f181-4ccf-8356-9aeb00c7c3af · outbound

This paper cites Language to Rewards for Robotic Skill Synthesis.

LLM-as-a-Verifier: A General-Purpose Verification Framework Language to Rewards for Robotic Skill Synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:823f7a2ed08228f14b68ef40e622a62de8dafff2dabef0abebe69db92956b502

Observation 452ca998-7168-4afb-8c1e-428707d1d97f · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

LLM-as-a-Verifier: A General-Purpose Verification Framework Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:666d95982de9baa4f24f0f004bb064cb6fe17aedad7ff25369f78c6c21d2e5ef

Observation 0172d717-a7b0-4903-b8a0-27daae6f082d · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:773d45baba92961ee78b00885c43ec20eac209dd9f1a087cfbc01d8c63c87c93

Observation 8f8f2769-e3e1-425b-b199-90a7a0b636e9 · outbound

This paper cites Lim, Jesse Thomason, Erdem Biyik, and Jesse Zhang.

LLM-as-a-Verifier: A General-Purpose Verification Framework Lim, Jesse Thomason, Erdem Biyik, and Jesse Zhang

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:d0206a979571b36d97dfa49ed73275b012a068233bcca9534110cdcb5d5ef237

Observation 7362ffb4-624c-4e66-8651-8578f6e79d9c · outbound

This paper cites SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation.

LLM-as-a-Verifier: A General-Purpose Verification Framework SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:7ee69b983f70b3cae26cd63dae3ec9a3a570f6914a7940a14626e1d31f17d375

Observation d644c3bd-b56f-4e03-93a7-2505c60d8a8c · outbound

This paper cites Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance.

LLM-as-a-Verifier: A General-Purpose Verification Framework Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:52feafcd817b0d2aad3774f7636b6ebaa17be1c837333a7fa3ba7398e33db3f4

Observation e3eb01b9-be0d-4184-b697-e308253d8c17 · outbound

This paper cites Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling.

LLM-as-a-Verifier: A General-Purpose Verification Framework Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:ad890d5f1ec67b864069c096449725dcc80d16a97bb753656382dc88b821df4d

Observation f30661cf-cd98-40ae-8a34-d91f52a39059 · outbound

This paper cites RoboMonkey: Scaling test-time sampling and verification for vision-language- action models.

LLM-as-a-Verifier: A General-Purpose Verification Framework RoboMonkey: Scaling test-time sampling and verification for vision-language- action models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:655bb28d0e6b35c5a725c96f239eae19ec7bbbc0212888b91269d55a8750e763

Observation b32664f1-d036-4651-88f2-b39bacacde24 · outbound

This paper cites Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress.

LLM-as-a-Verifier: A General-Purpose Verification Framework Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:77aaedebfe63cdc96be0dd9b618aa41ad43cadd77d445a334c559d380ed26687

Observation 29bdb0e4-9944-4d43-994c-52f6a0ad8a66 · outbound

This paper cites Scaling verification can be more effective than scaling policy learning for vision-language- action alignment, 2026.

LLM-as-a-Verifier: A General-Purpose Verification Framework Scaling verification can be more effective than scaling policy learning for vision-language- action alignment, 2026

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:823132cad158361331d8da32b3e8274a127ee5d0937c948cebeb6dbb9cb97eb6

Observation 2fdc84d5-7da1-49ee-bff1-d66f2c67b25a · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

LLM-as-a-Verifier: A General-Purpose Verification Framework FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:68fc32c0862eb0c8303a5982add6c47767cbe596bfba48836fb8b88521225fc5

Observation 726055b4-c0ff-4fc2-b31b-00150c257451 · outbound

This paper cites QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization.

LLM-as-a-Verifier: A General-Purpose Verification Framework QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:0259c57b160dbd36f382ca920c66e36f09962ad371f2ebabc6e22766c2283cde

Observation d09b7fde-c598-4851-98a9-1d61244a3156 · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

LLM-as-a-Verifier: A General-Purpose Verification Framework SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:95af6be9aff39e0fef19b57ec6a5bbc6fbc9fbaa45d2ec2e6a5297d2599038f1

Observation 737ecf2c-6276-4a63-a36c-64a89a504146 · outbound

This paper cites World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry.

LLM-as-a-Verifier: A General-Purpose Verification Framework World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:dd00e0538ae9216d176c89a878fa941b38cafc4feca33fe4e8b570ca868ba713

Observation 295e4766-fdcb-47ea-9d6d-291d7fd652fc · outbound

This paper cites SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation.

LLM-as-a-Verifier: A General-Purpose Verification Framework SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

Reference 95

Resolution
malformed identifier
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:ade764439494eba3612bb4d2a7d5df5ce011745ed4dd8127cfa2860c627a7362

Pith citing papers

Observation d1f2b67b-4c8d-4134-a93a-2b055208e62b · inbound

Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration cites this paper.

Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration LLM-as-a-Verifier: A General-Purpose Verification Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T17:46:56.897385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:46:56.897385Z digest=sha256:71192d58aec3c21f8cd345e7ed3a4a358a707ff47003ded33c4c3e99d1eb92dd

Observation 5635ad34-a720-4065-b69f-b132099a4386 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning LLM-as-a-Verifier: A General-Purpose Verification Framework

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.512215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.512215Z digest=sha256:a2217a7658843b41052df82c1ab38d981dfe4c5e96d50134f4285cf04d19271f

Observation 324234cf-96c5-40a6-9c77-9101aa57a1d3 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning LLM-as-a-Verifier: A General-Purpose Verification Framework

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:25.009152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:25.009152Z digest=sha256:0ae8d966835faacb3e24efea98962342680f3d034d6070828d4e4e799ba8e664

Observation a8dd1fcd-b443-4331-a919-1f7c28ec0438 · inbound

When Policies Change Probabilities: Modular Decision-Making for LLM Code Review cites this paper.

When Policies Change Probabilities: Modular Decision-Making for LLM Code Review LLM-as-a-Verifier: A General-Purpose Verification Framework

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:12:15.996343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T00:12:15.578147Z digest=sha256:c06f988f75992e5ef39d1606131ab9c510c5d045f190f97ef1ae6229e2ce367d

Observation e511c256-8f9b-42bf-8f51-b1c08c10a7fd · inbound

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments cites this paper.

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments LLM-as-a-Verifier: A General-Purpose Verification Framework

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-10T04:59:12.973916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:59:12.973916Z digest=sha256:298f0f9589f68b3f55624328b0f7079c510499eccf5c6e28d81a5a8f47ebd851