Pith. sign in

Paper Citation Record · LEDGER

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

As of 17 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 10 inbound Pith citation observations for arXiv:2504.15253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15253 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.636506Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:49:05.104352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.244210Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved67
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 451c1f70-5522-4b0f-9ee9-fc2aaf2dd917 · outbound

This paper cites write newline.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.300764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.300764Z digest=sha256:bda4160e3629baa9e797f66ca66c3cbc6613f178dad48063556ca9e347eafe45

Observation 1e018f6e-37b3-4206-a847-e0be8b76d140 · outbound

This paper cites Program Synthesis with Large Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.306561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.306561Z digest=sha256:2208499d4543da753eda9dc04d7f6e5791eb25bc0240a209b9e8e289b20b15db

Observation d3fd40bc-5e7f-479e-adc7-90325d5e8077 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Graph of thoughts: Solving elaborate problems with large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.855145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.311394Z digest=sha256:29c8de6ba43149eeae3788ab11941318eab33aaf0cb25e88782b147efecd5146

Observation 8d9a5256-1a74-4b01-89b0-92406c8ab486 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.315974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.315974Z digest=sha256:bb803b6e22bfe5ce77a7ce9684f0094fe96e300f3d66356355ba4d0e50e903ae

Observation 6734b43b-77e0-4d0f-87b0-b27b138c6bd4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.320892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.320892Z digest=sha256:05daf2b6acaf2070af31955450effc8634328c5db03150c63e52e66a5b0b8999

Observation 0e9e4198-24f9-4b71-b9e6-4a3fb9f7094b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.325919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.325919Z digest=sha256:09b57a4dd0b3c400e29221945927f9d10a93d573de81198a6b96ed6cb43330bc

Observation 5382001d-e052-4af1-a101-b649e9e57838 · outbound

This paper cites Process Supervision-Guided Policy Optimization for Code Generation.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Process Supervision-Guided Policy Optimization for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.330823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.330823Z digest=sha256:45d0bf7917f15801c685047ca42c4a242500f8eccdfa1393a364ad66c435e2ec

Observation 46968a37-e714-4b95-81dd-993715acf818 · outbound

This paper cites GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.336123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.336123Z digest=sha256:ca17c7669b6e12260ffba9da8ba9bb2349cfb8917575a2bbf9a086c1b709674d

Observation 8e0849b1-8c21-4331-84d1-6e3a8d375610 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.341100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.341100Z digest=sha256:3e9fee11058108cc4ce2ffd76f8a1be51faa7a201cfdd61e638f4dbd57db5b8b

Observation e6daa37d-52c6-474b-acd6-ec7fa12234bb · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.345382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.345382Z digest=sha256:cd44a34b44c3878411d5c80485ec99dddb17e51b864c560715e76deea77c169f

Observation 85ac88c3-18a2-4a08-8e08-2f353d058c93 · outbound

This paper cites X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.839906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.349928Z digest=sha256:e4c87bca437d4f72aa49918ef1da5ac99da8d3d2ee13cfa7fb2c8ae8e9ed3251

Observation 7675ace2-b13c-44df-9582-dab965298b83 · outbound

This paper cites Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.354266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.354266Z digest=sha256:552a0e45de00ee15f43425718c3431e9f67d6fb4f7aee639078c65a1559336b3

Observation 245b1fac-f3ff-4cb2-bdb8-45ead85e4fab · outbound

This paper cites How to Evaluate Reward Models for RLHF.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators How to Evaluate Reward Models for RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.358755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.358755Z digest=sha256:a369d9381da57a0c41f39e0d0fc1336f03a3601ecb4d78caa8d67fa6f789f993

Observation 718aeae7-a6f7-48db-bb6b-75809036aa84 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.363247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.363247Z digest=sha256:8e5205a4221595755710dea9b8b42cd1d2b93585c574ccf0bc9147cb6a414f8b

Observation 94484cd0-cd2e-4ee2-96f2-51cbea4c2f9d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.367809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.367809Z digest=sha256:d779b3fa95cb51086011b40008a6b52f1531a1089aa9f435b01cebeb31a0c419

Observation 50da763f-3f6e-4d99-ae73-32b655d2f706 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training Compute-Optimal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.372411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.372411Z digest=sha256:1fd43b3c93a5637774d242d14c35820d84d09940779ec1fea368f98b95ad3afc

Observation b2fbcddc-be2e-4b00-9041-1ceb52025062 · outbound

This paper cites Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.377306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.377306Z digest=sha256:bdc5107af95d3b994e4c6c66da9049b04461d64701a48ebc69175c4909071870

Observation e5274ca1-0998-49b3-bad0-771f4fb2b6e2 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Large Language Models Cannot Self-Correct Reasoning Yet

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.381904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.381904Z digest=sha256:910195db33a8bc2922655e2b269e24e07caf1962a325ccdd12c5c0f306171341

Observation 16cfd505-a884-41f9-bfed-d064cdbc8927 · outbound

This paper cites OpenAI o1 System Card.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.386368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.386368Z digest=sha256:5a77ca419a3737e8b1ddf3f5040fbf05abbf923a86d2df74d0042f30ba204157

Observation 98edb3db-bb0c-4ce8-8863-150c35e6c302 · outbound

This paper cites A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.390529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.390529Z digest=sha256:ba198aa0ee5035b5f5ac2e25fe0b106f1311667106eb4629dcba1aa813976481

Observation cd580d81-d1da-463a-8886-9c18146aec37 · outbound

This paper cites Scaling Laws for Neural Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Scaling Laws for Neural Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.395112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.395112Z digest=sha256:cd00bd9183b99b1f4fce9c0503e72b03d13ba15cea441476f7b972a943fe4d7d

Observation 741aaaec-75c1-418c-ac2c-524de9216fc7 · outbound

This paper cites X., Li, M., Qin, C., Wang, P., Savarese, S., et al.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators X., Li, M., Qin, C., Wang, P., Savarese, S., et al

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.399161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.399161Z digest=sha256:dbf7486ce42e4ec82667425ff605885e592b2ac33938259b2546e0bc6d7cb031

Observation 478c77ed-6022-478c-98c7-a2628de02832 · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus: Inducing fine-grained evaluation capability in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.403828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.403828Z digest=sha256:e4585191975bab7113311294d16e900fa7d7803225a7fd956873304e7ff64677

Observation 3598cd01-4af1-4f62-99a6-b6773f0a2feb · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.408060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.408060Z digest=sha256:b50ea214b7d6bb11c0eb6a6b8c2860593bb2c711032d96f43336e6058afce743

Observation f0f36a76-4d25-4b65-a4a3-121f6a7bd1fe · outbound

This paper cites S., Reid, M., Matsuo, Y., and Iwasawa, Y.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Reid, M., Matsuo, Y., and Iwasawa, Y

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.412740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.412740Z digest=sha256:d4d0f3fb5855de1da580a5a24a2e618ceeee59a713c9c638fc50422f2f1857f5

Observation 06c7b997-4740-48d4-a1d8-939997b73ae0 · outbound

This paper cites H., Gonzalez, J.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators H., Gonzalez, J

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.416910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.416910Z digest=sha256:aae17b9bbf65347b28d9ec3d62f6e39914724940412ffa835eff049ad5a8eb85

Observation a8d4e9df-5494-4c13-b0ba-be468bf2a113 · outbound

This paper cites Math-Verify: Math Verification Library , 2025.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Math-Verify: Math Verification Library , 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.798859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.421111Z digest=sha256:9c466e04ecd510ae12f37ef1f40604881047a374eb181bff2bf821e73704e73f

Observation aaa7a588-d331-4ce9-9af7-3d02a7a0495b · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RewardBench: Evaluating Reward Models for Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.425258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.425258Z digest=sha256:cc45620b4642a16e4fec2bbd4266cd79d3d40f4990f729b1724dbf354d4e47df

Observation 24a687f5-59d2-46b2-a294-3d73cfbc7814 · outbound

This paper cites CriticEval: Evaluating Large Language Model as Critic.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticEval: Evaluating Large Language Model as Critic

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.429585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.429585Z digest=sha256:25fdafc7762c3beed74e8a9ad68f797cd82e33b398173ca8ca2f564400687750

Observation 0c69488b-b291-4319-ab87-4ff79f3be329 · outbound

This paper cites Generative Judge for Evaluating Alignment.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Generative Judge for Evaluating Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.433729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.433729Z digest=sha256:d790e05f7dc5b089e9a61f182de95bfaedfa662dde1f0c8127cc60a9c661cb0c

Observation f58363e7-5549-46c9-83a7-01634021b3ce · outbound

This paper cites Let's Verify Step by Step.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Let's Verify Step by Step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.438021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.438021Z digest=sha256:cc6e8cda8035a95fbe45ae2e219ce842ee30bb75001a2d3719f7575d9d0dd4f3

Observation e31cffcc-7df2-4230-b851-9b85a6631491 · outbound

This paper cites CriticBench: Benchmarking LLMs for Critique-Correct Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.442371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.442371Z digest=sha256:edecb484325c22ff04d816c9a1a69fbb50d62a350ab2a20675ae598b32b38579

Observation 55611e9e-14d5-4da4-ba8f-e0885cb0dfe2 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.448043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.448043Z digest=sha256:7a631a1fa9b78d425232338842cd36ff40970a186cb8deedb772873120aebaa5

Observation 496ae6bc-ffce-44ee-a05d-039d8c7124ef · outbound

This paper cites S., Wang, Y., and Zhang, L.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Wang, Y., and Zhang, L

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.452426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.452426Z digest=sha256:ff57257d78da66e38e317c78cfb26880fefd5e0ee0153c892e56d304cca3d2a9

Observation 643549e0-a18c-401e-8fba-9d944bfb4a08 · outbound

This paper cites Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.456653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.456653Z digest=sha256:075764f6e7035ddc4cf720de21a023a6c50e21d899a99fad91ada0de9065c62a

Observation 78b95e4f-913e-434f-af6e-f7fc95ba8fb5 · outbound

This paper cites RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.460702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.460702Z digest=sha256:62d132e8dafba6508204be3b4ef14a6df0fa17f27e0fbcb0cd8d202f1777459c

Observation 2567b4ed-d673-475c-82dc-496b3fe088e1 · outbound

This paper cites PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.465051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.465051Z digest=sha256:42b54ea1cf90aa2d7e2b85096c933e35ffbf17a4a12a7ccbd2ed24ad0906d340

Observation 0c122cdd-c3a1-4556-844c-e53606f84df6 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.469382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.469382Z digest=sha256:d657f789db282d505a88098c8d4d407b61537b921b28f3f38f428bedb0442367

Observation 3d1c0f5d-5c49-4186-89e1-780c76352aa5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-refine: Iterative refinement with self-feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.473889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.473889Z digest=sha256:3737ba9e803e8aac50cce4146184ff0634345b30898e9cdb36284bfd4251d0cd

Observation 977c2e83-8fa3-429a-a89e-285fad4c092e · outbound

This paper cites CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.478015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.478015Z digest=sha256:afd6cb7afd349cdf477f104d1d49884baa9b2c6b141bb8c4309dd1466eddf517

Observation 7191b780-1b64-44d8-9b56-ba9b3a3e6bdd · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.482766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.482766Z digest=sha256:375b03a11b89f7b3bc482098c8ae6bc10fe3a95a3bda8c951ca88772406314f7

Observation 82918c6a-4387-4887-a60e-c49d60dbb4db · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training language models to follow instructions with human feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.487338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.487338Z digest=sha256:e7cd9f0432d9f57ced27a27179e7e786a50c7d94bddf1219fe0fb97d35a8bdb8

Observation c8014eb0-9e0e-4c9d-b102-143e4b4e20af · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.491812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.491812Z digest=sha256:e1b3331e0298d5f01da01c6808191ba5d0f7d189617586d7e78237ef28bd07ef

Observation 7d945d56-b8ce-42cc-bd2d-e87388d2bd76 · outbound

This paper cites S., Franklin, M., Vidgen, B., Singh, A., Kiela, D., and Mehri, S.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Franklin, M., Vidgen, B., Singh, A., Kiela, D., and Mehri, S

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.496376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.496376Z digest=sha256:5f33a4ab793c039ff0cd0a69264e260bbbeeb6b9feaad1494181a59feb4c34d4

Observation de9f9b56-54bd-4a22-83e2-e57183e165ac · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-critiquing models for assisting human evaluators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.500649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.500649Z digest=sha256:8dc2f9341a4f9ede1e1bb8d8b83c062c32eabe8f153c48fde2d2fb3abe7a2152

Observation c7149a45-f235-4072-ac3d-1408ab792ee6 · outbound

This paper cites B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.505284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.505284Z digest=sha256:a0fff8389b80741aed81062b3e0ae0783fb3172d32bfe76b8f8c0d6b1d6900cd

Observation 84c7b553-33c8-4038-b7ca-83bf8d5f3c8c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.509484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.509484Z digest=sha256:d705b8585ddd0dda84d31ce56c8e849f295b23bea81e4dc8dac687e8aee269e2

Observation 60870029-d586-48a4-9a55-6e9654e0b148 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Reflexion: Language agents with verbal reinforcement learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.513992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.513992Z digest=sha256:89307f2a81761ba831d46300c7768bed8500522b228aa1d6da89732fbc63090e

Observation a1f1f110-c69c-4b07-956a-f2a978274924 · outbound

This paper cites Y., Zeng, L., and Liu, Y.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Y., Zeng, L., and Liu, Y

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.741694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.518677Z digest=sha256:1c19eef70e5ba183f7074565c8ef080d4a7bf2bdbae54c5851545640cd694862

Observation 6859ef88-2e65-4791-8061-39c3d1502403 · outbound

This paper cites The art of llm refinement: Ask, refine, and trust.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The art of llm refinement: Ask, refine, and trust

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.728380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.522865Z digest=sha256:37bc887c1a2ee6f7c24b19dc23256c35eb1a67cdf1d010a102de494f5fb62ee1

Observation 3dece122-5448-4b6d-8dbc-e0a8d63d92ba · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.527127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.527127Z digest=sha256:f9a7416c114ede243fad615832e45700752c584e0e416019a8a641738c97e4d0

Observation 536088e7-b448-45e6-9de9-e38a6752366d · outbound

This paper cites GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.531866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.531866Z digest=sha256:ca36fb2e61f38db33bf5542e609c707fb381c5e04291fe89eac4707d698d4a69

Observation e18be681-439a-4905-a4ba-f8122d176cc5 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.536375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.536375Z digest=sha256:c6aa5ed386d65d9490679870fe4528a363e47ade91a433996a455f6acaa153b5

Observation 954c0c8f-311b-4be5-9824-a8c4696ce09c · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.540998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.540998Z digest=sha256:c48e10f36d820fc0dce9bb510dedc9d046bee91332d3bd7323655cdf4de04ca0

Observation 4396bc91-83a8-44e2-88cc-fb980e50cc79 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.545529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.545529Z digest=sha256:54f827929c3b4528a1e8141b957321e64c2b9de540d7d3fe904b4cd747f4f6f5

Observation 5e6632f4-8e24-4e34-b0ad-ea2e400f7683 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.550130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.550130Z digest=sha256:2d5722693707f9457c6fc4a086a2965b690c05ff0351a729c57492789d89ddb3

Observation c5ef18c6-baba-4b8d-9333-4ecf6eb80d1b · outbound

This paper cites Direct Judgement Preference Optimization.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Direct Judgement Preference Optimization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.555545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.555545Z digest=sha256:a6c359a693a2fcb46a4a80054e97fb9c4fe1e60ac321a8006f1fef9154ff7ece

Observation 701a70b4-f8a6-47a9-bc3b-267b2ab298cb · outbound

This paper cites Self-Taught Evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-Taught Evaluators

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.560109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.560109Z digest=sha256:a5b169668c240043f70423d1e3fc99d5789edca0e2247ceba5bb4bcf498cd394

Observation 44ad93b7-01c2-4723-b5da-506d0a383821 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.564625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.564625Z digest=sha256:8c0a55c39e4528bdd521753cfbcdb427226a1008dcdf01b6af276f788babba52

Observation f43c3f82-067e-49ec-8574-7cbffc4cc8e9 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.714084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.569347Z digest=sha256:2e1c8952e3cb351d977878bfc805be03f8d5c2fd89863c617efa091db7a7184f

Observation 2ffd2a87-60c8-4488-b6b3-6768ed2f5116 · outbound

This paper cites V., Zhou, D., et al.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators V., Zhou, D., et al

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.573865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.573865Z digest=sha256:003e97426e4d64332ffda000458db68a8b661d32b79553b783735baa587bc76c

Observation 7690cba8-e0a8-4132-9453-0385b6bbbd3d · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Qwen2.5 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.578079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.578079Z digest=sha256:91d67d0257f2317d7b7b7c17e396c823fa85486c09d83435d3bbe976d953f5eb

Observation 0d165790-bc5c-4919-acc8-6e354c6a6d38 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Tree of thoughts: Deliberate problem solving with large language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.582331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.582331Z digest=sha256:b154c0cfcb10bbb6becb4b3962b165ce59f7208e08a73d07f1390432d3ce3fac

Observation 5b3dfb1a-361f-4133-8fb0-96b2811fb289 · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.586506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.586506Z digest=sha256:5e7dc0981436d85956c26ced334e0dda47f7846e4a941e1b00eee53011f52b17

Observation 53dcbc1c-be50-4ba6-951b-5dc1a141dbc5 · outbound

This paper cites Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.681184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.591291Z digest=sha256:4a5d599b9411cab36f382bec029e0d202459da92688ceb903e95cf54cf85081a

Observation eb76a554-c718-4ac2-857d-9205a4eaa57e · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Evaluating Large Language Models at Evaluating Instruction Following

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.595367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.595367Z digest=sha256:506edf4026b82695970b6056903f1c600893ef5b904b149b76cc0a16c9a05f16

Observation 95280075-68d4-4a70-8c31-ed4e79f19255 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.600017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.600017Z digest=sha256:a40aaf0489313e415dd04d66a1aec0b177064f6e5e26c59d3d5fd797439df7f1

Observation 54fb66db-55dc-43d6-a631-c2d3b0562de9 · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.604669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.604669Z digest=sha256:9c6666549ef212f9dd16fb2302fae0edd33ad6dc656018469cfea83d8f6739f5

Observation 90a44f28-9f90-41be-ae81-addd9d3b8ecc · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.613435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.613435Z digest=sha256:feaa7de37acad5abc0a796a70b1bd42df443a699392adff75d2387df0d9c0132

Observation a4613a37-9cde-4a00-8718-f664a494ef58 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.617714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.617714Z digest=sha256:fbb90f801a3a2428650c43d3ff76565ca6268565d84c6a597b645136664db153

Observation a5994195-4f4c-4b6c-966e-692c32e0dd98 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.622552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.622552Z digest=sha256:dbc34a11888d53fbe5ffaa1dc1462fd9313d1ec016cc4a56bdb6256d7cd9f61f

Observation c567d069-316d-455a-9c12-4d8d5de7d696 · outbound

This paper cites RMB: Comprehensively Benchmarking Reward Models in LLM Alignment.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.626881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.626881Z digest=sha256:9c8bb65fc70e01394bd9e126cf0c19292ad4d0cd13621d626b3901587ef149c2

Observation 72a44a73-3fdc-4843-bbf8-c12095820ac7 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Instruction-Following Evaluation for Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.631721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.631721Z digest=sha256:e4343e8f6127c7f1e3cf5a51b9584832debefef140c5d88d6a8d265c88113f5d

Observation 3c1b5ab9-6c50-4923-8389-96ab9666bb70 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.636506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.636506Z digest=sha256:24bb0ddf2a19ccc18f4a4ee3de0f0386dcf2d55eeaa7a7a2113bff1bf1cf4246

Pith citing papers

Observation 9718b30f-8129-4090-b209-9e0b43a91073 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T21:56:01.605280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:56:01.605280Z digest=sha256:eaf22862765db12f89935999058e4abd2d1c1ab3702b8b8619f80f30708fdde9

Observation aa498442-b8c9-46d0-9379-96303932e058 · inbound

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models cites this paper.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:56.146103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:56.146103Z digest=sha256:503cf765ac16fbc4187d5bfe50ec25c82efe778ad2e5486d31675ecc9fb6e861

Observation d3531d6c-a718-48ad-b4dd-4bfeb6ed5f0a · inbound

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning cites this paper.

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:47.902627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:47.902627Z digest=sha256:b0d1f167cfd99a3d776d2059502b8989b8d80e046864e6930642217146a3c036

Observation 2052e7dc-8e34-4643-960f-36bb1c65a4ba · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.588950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:ef24b3074713437a090e14a5d110101eaa166a8c10299919b68c0f6671e90a83

Observation 142f02d3-7d4c-4ff0-abbd-df389318e56e · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:14.991198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:706c3d04275bb524e4fbc2741df7b81538874b80ae6f4f026a36c76660e8b277

Observation d2d5591c-75b5-4c20-a51b-0e8e55a72838 · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.679091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:0c46337239d3082cfaf35c45c9244a5678916058a5c1d5c510ba5ad27dcd900f

Observation df095008-aa51-4e87-9f08-e3532340a4ae · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:49:38.246250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:038077a093cc939619a5d659d109ecd3c79a0c3562b6910993cd20348da0c202

Observation d44d8c1f-a5a3-4d80-88a5-19b8ac73ceec · inbound

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification cites this paper.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T02:43:21.225324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T02:43:21.225324Z digest=sha256:2ab849e473c915d55d35520170fa7871aaa33d97b4af1d579b9fbb9b3734974d

Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · inbound

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning cites this paper.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.038013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.038013Z digest=sha256:60ed57c4ff28c9722eb2854b80cacd05b56dee2927c64c5557e168c370d83d0e

Observation 82ca2942-d1db-4f9d-aa2d-2b2159e19bfe · inbound

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility cites this paper.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:05.104352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:05.104352Z digest=sha256:630879e09b016c8a72e8a77dd87f70cfb4631b13e14648366a1dbea9f65f4eab