Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.636506Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 10 inbound Pith citation observations for arXiv:2504.15253.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.636506Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:49:05.104352Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:49:38.244210Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 451c1f70-5522-4b0f-9ee9-fc2aaf2dd917 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e018f6e-37b3-4206-a847-e0be8b76d140 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3fd40bc-5e7f-479e-adc7-90325d5e8077 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Graph of thoughts: Solving elaborate problems with large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8d9a5256-1a74-4b01-89b0-92406c8ab486 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6734b43b-77e0-4d0f-87b0-b27b138c6bd4 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9e4198-24f9-4b71-b9e6-4a3fb9f7094b · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5382001d-e052-4af1-a101-b649e9e57838 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Process Supervision-Guided Policy Optimization for Code Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46968a37-e714-4b95-81dd-993715acf818 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0849b1-8c21-4331-84d1-6e3a8d375610 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6daa37d-52c6-474b-acd6-ec7fa12234bb · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ac88c3-18a2-4a08-8e08-2f353d058c93 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7675ace2-b13c-44df-9582-dab965298b83 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245b1fac-f3ff-4cb2-bdb8-45ead85e4fab · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators How to Evaluate Reward Models for RLHF
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718aeae7-a6f7-48db-bb6b-75809036aa84 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94484cd0-cd2e-4ee2-96f2-51cbea4c2f9d · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Measuring Mathematical Problem Solving With the MATH Dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50da763f-3f6e-4d99-ae73-32b655d2f706 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training Compute-Optimal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2fbcddc-be2e-4b00-9041-1ceb52025062 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5274ca1-0998-49b3-bad0-771f4fb2b6e2 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Large Language Models Cannot Self-Correct Reasoning Yet
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16cfd505-a884-41f9-bfed-d064cdbc8927 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OpenAI o1 System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98edb3db-bb0c-4ce8-8863-150c35e6c302 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd580d81-d1da-463a-8886-9c18146aec37 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Scaling Laws for Neural Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741aaaec-75c1-418c-ac2c-524de9216fc7 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators X., Li, M., Qin, C., Wang, P., Savarese, S., et al
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 478c77ed-6022-478c-98c7-a2628de02832 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus: Inducing fine-grained evaluation capability in language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3598cd01-4af1-4f62-99a6-b6773f0a2feb · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f36a76-4d25-4b65-a4a3-121f6a7bd1fe · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Reid, M., Matsuo, Y., and Iwasawa, Y
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06c7b997-4740-48d4-a1d8-939997b73ae0 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators H., Gonzalez, J
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d4e9df-5494-4c13-b0ba-be468bf2a113 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Math-Verify: Math Verification Library , 2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aaa7a588-d331-4ce9-9af7-3d02a7a0495b · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RewardBench: Evaluating Reward Models for Language Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24a687f5-59d2-46b2-a294-3d73cfbc7814 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticEval: Evaluating Large Language Model as Critic
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c69488b-b291-4319-ab87-4ff79f3be329 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Generative Judge for Evaluating Alignment
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f58363e7-5549-46c9-83a7-01634021b3ce · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Let's Verify Step by Step
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31cffcc-7df2-4230-b851-9b85a6631491 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55611e9e-14d5-4da4-ba8f-e0885cb0dfe2 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496ae6bc-ffce-44ee-a05d-039d8c7124ef · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Wang, Y., and Zhang, L
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 643549e0-a18c-401e-8fba-9d944bfb4a08 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b95e4f-913e-434f-af6e-f7fc95ba8fb5 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2567b4ed-d673-475c-82dc-496b3fe088e1 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c122cdd-c3a1-4556-844c-e53606f84df6 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1c0f5d-5c49-4186-89e1-780c76352aa5 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-refine: Iterative refinement with self-feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977c2e83-8fa3-429a-a89e-285fad4c092e · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7191b780-1b64-44d8-9b56-ba9b3a3e6bdd · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Show Your Work: Scratchpads for Intermediate Computation with Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82918c6a-4387-4887-a60e-c49d60dbb4db · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training language models to follow instructions with human feedback
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8014eb0-9e0e-4c9d-b102-143e4b4e20af · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d945d56-b8ce-42cc-bd2d-e87388d2bd76 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Franklin, M., Vidgen, B., Singh, A., Kiela, D., and Mehri, S
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9f9b56-54bd-4a22-83e2-e57183e165ac · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-critiquing models for assisting human evaluators
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7149a45-f235-4072-ac3d-1408ab792ee6 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c7b553-33c8-4038-b7ca-83bf8d5f3c8c · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60870029-d586-48a4-9a55-6e9654e0b148 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Reflexion: Language agents with verbal reinforcement learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f1f110-c69c-4b07-956a-f2a978274924 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Y., Zeng, L., and Liu, Y
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6859ef88-2e65-4791-8061-39c3d1502403 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The art of llm refinement: Ask, refine, and trust
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3dece122-5448-4b6d-8dbc-e0a8d63d92ba · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536088e7-b448-45e6-9de9-e38a6752366d · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18be681-439a-4905-a4ba-f8122d176cc5 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954c0c8f-311b-4be5-9824-a8c4696ce09c · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4396bc91-83a8-44e2-88cc-fb980e50cc79 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e6632f4-8e24-4e34-b0ad-ea2e400f7683 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ef18c6-baba-4b8d-9333-4ecf6eb80d1b · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Direct Judgement Preference Optimization
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701a70b4-f8a6-47a9-bc3b-267b2ab298cb · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-Taught Evaluators
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ad93b7-01c2-4723-b5da-506d0a383821 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43c3f82-067e-49ec-8574-7cbffc4cc8e9 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ffd2a87-60c8-4488-b6b3-6768ed2f5116 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators V., Zhou, D., et al
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7690cba8-e0a8-4132-9453-0385b6bbbd3d · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Qwen2.5 Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d165790-bc5c-4919-acc8-6e354c6a6d38 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Tree of thoughts: Deliberate problem solving with large language models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3dfb1a-361f-4133-8fb0-96b2811fb289 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53dcbc1c-be50-4ba6-951b-5dc1a141dbc5 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eb76a554-c718-4ac2-857d-9205a4eaa57e · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Evaluating Large Language Models at Evaluating Instruction Following
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95280075-68d4-4a70-8c31-ed4e79f19255 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fb66db-55dc-43d6-a631-c2d3b0562de9 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Small Language Models Need Strong Verifiers to Self-Correct Reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a44f28-9f90-41be-ae81-addd9d3b8ecc · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4613a37-9cde-4a00-8718-f664a494ef58 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5994195-4f4c-4b6c-966e-692c32e0dd98 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c567d069-316d-455a-9c12-4d8d5de7d696 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a44a73-3fdc-4843-bbf8-c12095820ac7 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Instruction-Following Evaluation for Large Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1b5ab9-6c50-4923-8389-96ab9666bb70 · outbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9718b30f-8129-4090-b209-9e0b43a91073 · inbound
Can You Trick the Grader? Adversarial Persuasion of LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa498442-b8c9-46d0-9379-96303932e058 · inbound
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3531d6c-a718-48ad-b4dd-4bfeb6ed5f0a · inbound
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2052e7dc-8e34-4643-960f-36bb1c65a4ba · inbound
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 142f02d3-7d4c-4ff0-abbd-df389318e56e · inbound
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d2d5591c-75b5-4c20-a51b-0e8e55a72838 · inbound
Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df095008-aa51-4e87-9f08-e3532340a4ae · inbound
Counsel: A Meta-Evaluation Dataset for Agentic Tasks Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d44d8c1f-a5a3-4d80-88a5-19b8ac73ceec · inbound
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · inbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ca2942-d1db-4f9d-aa2d-2b2159e19bfe · inbound
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 151
Source-reported events for the cited work
Unavailable: canonical work link unavailable.