Pith. sign in

Paper Citation Record · LEDGER

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.01000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01000 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:43:01.149558Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact4
  • verified fuzzy22
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56765134-c1f1-45e3-8ad6-3b61f52374e3 · outbound

This paper cites Concrete Problems in AI Safety.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:55.654892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:55.654892Z digest=sha256:c81cdc0a6666a0a25674756c4a66fca269a969415175833d0e3e6884f822a7b1

Observation 07b5ba70-1fdd-497c-aafb-e2b2127b7b1f · outbound

This paper cites Program Synthesis with Large Language Models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:55.750021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:55.750021Z digest=sha256:390013d6a11c938adbd7ca61c8abaf65412d00bdbeb2f75f97c2281e83f1e87e

Observation 1b80d79f-183c-4525-94a0-aabd8181c5d5 · outbound

This paper cites Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:55.896502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:55.896502Z digest=sha256:bfad28f6df399b284a073b51bd9cf7097ad1e3b4f2e727d5ad66a9cd396d36a9

Observation c967abae-1ab3-4368-b4b6-5466d675666d · outbound

This paper cites Answer Matching Outperforms Multiple Choice for Language Model Evaluation.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.011737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.011737Z digest=sha256:f967a0927c568ad4623ea143670f5a1f0d1725ae5611d704eb518d1eade8620c

Observation 6850ce16-9810-4849-b0c8-5a01b6b1b022 · outbound

This paper cites CodeT: Code generation with generated tests.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets CodeT: Code generation with generated tests

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:06.627859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:56.147396Z digest=sha256:5e69661999544ee7512c7c7d2f3255f3afd3beeae5b6778759e5bb60258920a3

Observation a6f23190-75dd-4ea2-a758-ac7ab7893f39 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.255843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.255843Z digest=sha256:42d7fedf3367713f5a74d932d777495ae667f4053678183dcc6425bd36031947

Observation 80aa02aa-e708-4211-a92a-734accb60b35 · outbound

This paper cites Jimenez, John Yang, Kevin Liu, and Aleksander Madry.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Jimenez, John Yang, Kevin Liu, and Aleksander Madry

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:06.479453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:56.395232Z digest=sha256:004d39650b1681b09e2293e642d1dab86514235c566b76ce4ed05cf293f6e7ce

Observation f09a48a7-2071-4fff-9943-f16d5e6aaff7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.490998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.490998Z digest=sha256:4f5419991207dc560a114c6d36bae70e2a5e5e8a3ce87c7023fe0ccb202ba6d7

Observation 27ee833d-e729-49cc-ac06-0519fce26530 · outbound

This paper cites Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.617826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.617826Z digest=sha256:e54c50ece1a6e94adfdfa9440efb6422b776ed6463951139b1c633bb98787e9d

Observation 005b98e3-25d7-4240-a9d3-51f1bbc8ddb5 · outbound

This paper cites Rabe, Talia Ringer, and Yuriy Brun.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Rabe, Talia Ringer, and Yuriy Brun

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:06.329643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:56.724127Z digest=sha256:3f2ce8e870aa82ffdfa6c11a4d9263a9ceb76b1b50b32a2ea62978d324df5d2b

Observation e545781d-73f3-4ba4-ad35-7054ea7994a2 · outbound

This paper cites Scaling laws for reward model overoptimization.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Scaling laws for reward model overoptimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:06.203280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:56.793886Z digest=sha256:e4bd78383a9f66a18a6259e57f885ae06120203384ec0d44e538a404f724a196

Observation 801c6585-745f-478f-b774-0563d4008229 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets A Survey on LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.891870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.891870Z digest=sha256:05bb1195c17ad4830ca6ed012cad4d68b9dfbf6be8f43d94274cd45285bcd693

Observation 7c120526-1391-40cd-8620-153f1f432da0 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.941409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.941409Z digest=sha256:ba697f7322b7437c13a124a6614174e219e8a266a839e807d5fafc900e21b3f7

Observation dbdc489b-57d3-41ac-8f0c-cea0e5b2814b · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:57.018264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:57.018264Z digest=sha256:07acd3df478a2e0b7108fa07beaadb6b573fcb5d7b4c5a55aa5cc51d645f2e0f

Observation 6c547609-724c-468b-ac3d-2d5426bc755a · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Large language models cannot self-correct reasoning yet

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:06.049364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:57.177300Z digest=sha256:d3dccc2598a4d8f633e4bd1ca7bcb9e2bb6eb43bcdd339cdb59d00ae8cb82ddf

Observation a3212086-ea3f-4470-b589-c33e24fd66c1 · outbound

This paper cites Pitfalls of rule- and model-based verifiers: A case study on mathematical reasoning.arXiv preprint arXiv:2505.22203, 2025.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Pitfalls of rule- and model-based verifiers: A case study on mathematical reasoning.arXiv preprint arXiv:2505.22203, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:57.276071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:57.276071Z digest=sha256:6633595ccdfd73e7f60075109b64dc01e0ae00fdc8037d2062fd1d6e00f033cb

Observation c04fee6b-eb3d-492a-8552-fb573dc0c5a7 · outbound

This paper cites Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:57.335872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:57.335872Z digest=sha256:28b04d4cf4a5cafc91652b78317cf33daf2af0dc10516c0d40b02ffc45d01d08

Observation d4d63866-d6d2-445d-90d9-125f3954079a · outbound

This paper cites SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:57.387763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:57.387763Z digest=sha256:97acfe611651f26e8422784bb6fbdff6ab2057b4d7dec8b6664c8f865e33001a

Observation a89ff8da-017e-4a43-8066-a445deb8ca2d · outbound

This paper cites Language Models (Mostly) Know What They Know.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Language Models (Mostly) Know What They Know

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:57.450552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:57.450552Z digest=sha256:b2f68956c2dc57e2a5254eb2becacdd742e077860a4921f87d838a88fbd93a7b

Observation 31b5bf56-2455-40df-8c27-4ac6a37a2744 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:57.612250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:57.612250Z digest=sha256:072c73bae5978c30569384178682000958f69a312d8fd2ad72ec9c317bf72b5e

Observation 15d8bc4f-2ab0-47ad-996f-6442d3c29fbb · outbound

This paper cites Let’sverifystepbystep.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Let’sverifystepbystep

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:05.861746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:57.680434Z digest=sha256:7fdd0a71dac2fa393ec99ffa02cf2fde28a4e19245a63e08675c635b5f82531a

Observation 1041a728-0780-4d1b-9b1e-9fc7e41be5ef · outbound

This paper cites Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:05.690456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:57.779029Z digest=sha256:b047e14c7a58d500f7518e270f0ae22838e48ca17740e596d67214424aa8c197

Observation f45f340b-6e35-426b-a2c7-6c37c87c9d3c · outbound

This paper cites Rethinking Verification for LLM Code Generation: From Generation to Testing.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Rethinking Verification for LLM Code Generation: From Generation to Testing

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:43:02.301477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:57.915322Z digest=sha256:0d8ef1f5277f98152e1b0d7b8eb22765bd8850a4fe4108c84e85f7c99cdcccf2

Observation 261b4d69-af70-4247-8968-b61cb2c8d588 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-refine: Iterative refinement with self-feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:05.515856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:57.957263Z digest=sha256:c58149069a1a9ceae04de4d0cc6de71828280f42090e4c97c5fb9b32ab96a31c

Observation 5c55bc9d-56f5-495b-8c7a-f85f33a27070 · outbound

This paper cites an unresolved cited work.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:43:05.362211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.059997Z digest=sha256:55bff40ced85f390cbbbb14e7f99c5642c3a4b100476ed60208b191feb8e2b4c

Observation d4dedd60-c8fd-44a7-ae32-aa71ce43d2a6 · outbound

This paper cites What can we learn from collective human opinions on natural language inference data? InProceedings of EMNLP, 2020.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets What can we learn from collective human opinions on natural language inference data? InProceedings of EMNLP, 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:05.208678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.138372Z digest=sha256:7049db1e0ef79c0d50a19c6452938afefa4f6aaacf787d35bc589d4c4ec730d4

Observation 2ae0cc0b-53c8-4a2e-8a68-5bcdab945dd1 · outbound

This paper cites The effects of reward misspecification: Mapping and mitigating misaligned models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets The effects of reward misspecification: Mapping and mitigating misaligned models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:05.052619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.221825Z digest=sha256:1a732254b5d5ab1b510d0b5bb4f45ef2d87fd306d81d41102a8f4db41027a5ff

Observation a7d0c9a5-fc65-4dde-ac36-f927e43501aa · outbound

This paper cites Qwen2.5 Technical Report.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.258320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.258320Z digest=sha256:8ea97c63211dfc3044b4bb4fcdd867dc9f4a1ea9a1a7c9959fefa619402c60aa

Observation 9d5f5215-d663-4eaa-8ade-5d68212353fb · outbound

This paper cites Before the Model Learns the Bug:Fuzzing RLVR Verifiers.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Before the Model Learns the Bug:Fuzzing RLVR Verifiers

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:43:02.020167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.298232Z digest=sha256:206dae2d01ff5dc3cde17b7c447f59bf468f26c0660c9e14601752586dc82791

Observation 12ec88ad-73c2-45b3-82ff-215cc94ad8ef · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation.IEEE Transactions on Software Engineering, 2024.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets An empirical evaluation of using large language models for automated unit test generation.IEEE Transactions on Software Engineering, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:04.930480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.462050Z digest=sha256:17ea13d0905c760da820e01e45da36cbd8654ceeb0ef852344e6f17dca66526a

Observation 4d1e0d3c-c731-42f6-81d8-0aa460389574 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.610145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.610145Z digest=sha256:0cf53c0fc23e71d794ebb6def79215428b7e9d4836f0fe544f9d9377707fc317

Observation 59fe8eca-91b5-4b3b-b61f-47f671f12e7b · outbound

This paper cites an unresolved cited work.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.664880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.664880Z digest=sha256:eead00814a530b8f349419c89a3110620e2f2738878c7dbf1412cdfa96fe20c7

Observation 4a151dce-1fbe-4ea6-af68-2b1378e91128 · outbound

This paper cites Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.750607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.750607Z digest=sha256:65e56f72da42f76fabf905dc0b95d11533a16c1b24b012ef7e564b55c1634a07

Observation 0f457223-c64b-4f60-b8b0-d9575cbc3583 · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.819481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.819481Z digest=sha256:9cc4ef247a00ac512221ad5ab86036fea9308bfb9035cfda7424b5fc828f51c4

Observation 683e9c94-b538-447e-bfd9-f61edfaee109 · outbound

This paper cites Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:04.753162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.890270Z digest=sha256:92b5730744fc9bb48d200c7e0e26338554bdb927f36233c742d624a9c484673d

Observation 32f3b319-5db9-4391-9e00-d3e90967dfd7 · outbound

This paper cites Order matters: Sequence to sequence for sets.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Order matters: Sequence to sequence for sets

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:04.479794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:58.971370Z digest=sha256:acb369e6a788177192fb78776ddc536ddfeccab4e208749aa1ade34d5e3dea3a

Observation 290ed08f-de5f-440a-88ac-1ec852d72c44 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-consistency improves chain of thought reasoning in language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:04.312856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.054572Z digest=sha256:85e98d20c051226e3cf90e670b7b1e3a050460c39587f60641e662a20cf3dc0d

Observation 99421d6f-cc4f-4d6f-ac40-2c5440992a10 · outbound

This paper cites Con- sistency of a recurrent language model with respect to incomplete decoding.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Con- sistency of a recurrent language model with respect to incomplete decoding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:04.163427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.158105Z digest=sha256:1a7f5f85c58a37dffe7ee9536fc3547e9ac28ad5defe7d0bbe5e60867adf7010

Observation 2cc989b6-bc3c-4596-a067-a6cc37a92e28 · outbound

This paper cites Large language models are better reasoners with self-verification.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Large language models are better reasoners with self-verification

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:03.957082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.255283Z digest=sha256:b4aa560a750d6878d71f918c024622f777299cec58dc8331640387bd925a3675

Observation 3a972c1a-5c23-46e9-92af-72b9315f088c · outbound

This paper cites what it can create, it may not understand.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets what it can create, it may not understand

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:03.758630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.332880Z digest=sha256:8fe2b29c6705b4de215d4d910c50460cb01aa32b0102f3045fe536487f60399e

Observation 2df6edbf-9234-47ab-abbb-249b9208a7e8 · outbound

This paper cites Efficient Guided Generation for Large Language Models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Efficient Guided Generation for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:59.366613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:59.366613Z digest=sha256:27c1269893f3e03e726608da21165de25e4f713c461e20390df1f17083dbeaec

Observation 09c82dbb-8a4e-48cb-950a-43421840574c · outbound

This paper cites Jiang, Wenda Li, et al.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Jiang, Wenda Li, et al

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:03.588185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.427699Z digest=sha256:25e4b958e6c655c28b6ced0367e93c6c88c0f1b78d6bf23f2b529e162b99e3d6

Observation 18f93d2b-342f-4e6e-9566-7a552c9195ed · outbound

This paper cites Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:03.450792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.506928Z digest=sha256:2a6ebc58d6289db397b9f6ebbaace193b4bbb9dda27891692e90b42585d71354

Observation 0e153ff8-2c8e-4109-a9a0-2c5f7b589980 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:59.575499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:59.575499Z digest=sha256:b7f8d347e06503a3c8aaeaae3560dd62cae127bb886b734c3a2cfc505e86dc8e

Observation c690072c-e9ad-48de-9f5c-32e63bac7c6b · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Judging LLM-as-a-judge with MT-bench and chatbot arena

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:03.233303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:42:59.671625Z digest=sha256:19cbbab704bca8330c327897dfdcb684108b7cc960795fc0ebf79037bdca2087

Observation 047a85fc-af93-4817-91e6-e91dc3018dd5 · outbound

This paper cites RubricBench: Aligning model-generated rubrics with human standards.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets RubricBench: Aligning model-generated rubrics with human standards

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:59.770338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:59.770338Z digest=sha256:15784f7f86f4e0d654eac13d2296e8e90a52e10776af7e4b1846a5a8c9925b1f

Observation 1b7c10b3-f874-4fe8-8bb0-b59dfb1efac2 · outbound

This paper cites Self-Rewarding Language Models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-Rewarding Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:59.909123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:59.909123Z digest=sha256:9d22a91156d3eccb0b9868f0fa67af8d6b3fd16afa719dede91b392e7ac7c0d0

Observation ab446688-d7fe-443e-ab53-2f9e60002670 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:59.966639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:59.966639Z digest=sha256:96c925cad27fa329852144732ee5e993fd68af4100b3b0377f2f3318fb745343

Observation 12947eb9-15c3-41cd-ac6d-73160fddd018 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets LLM Critics Help Catch LLM Bugs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.039128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.039128Z digest=sha256:efe5aae9a8b8200587b285776d3cd2f6defebbfdf373329734be9b518002c0fc

Observation 796a9f74-3c3c-46c4-898b-aedb632ee5ab · outbound

This paper cites SWT-Bench: Test- ing and validating real-world bug-fixes with code agents.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets SWT-Bench: Test- ing and validating real-world bug-fixes with code agents

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:03.042207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:43:00.114410Z digest=sha256:011f602875a64abe6db0cc4a71ee0ac30fb96b164857b21e3e544777a5f24422

Observation 9a5417b1-96e7-432a-b22a-eeb1cf7271d2 · outbound

This paper cites an unresolved cited work.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:43:02.910960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:43:00.195636Z digest=sha256:bbae62c457594c57bce6d0b611d6755f16ca5d5e83f6269615fef1c02e99750c

Observation 7e78a78f-6223-4c63-9ea3-3b09e37e95c8 · outbound

This paper cites Tomayto, tomahto: Beyond token-level answer equivalence for question answering evaluation.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Tomayto, tomahto: Beyond token-level answer equivalence for question answering evaluation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:43:02.742530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:43:00.307939Z digest=sha256:9739c60b8ebf1a78ddade657db360f0678339ffa1b7716fb0fa8208e65451a6d

Observation 5cd6e2e8-5672-4516-a13f-9301b2efc1fc · outbound

This paper cites TOGLL: Correct and Strong Test Oracle Generation with LLMs.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TOGLL: Correct and Strong Test Oracle Generation with LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.408314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.408314Z digest=sha256:3995d11997b776765479f58799fb5f4b77a4596911c6163066a23ff0d2bf34a5

Observation ec88bab9-d4f6-4ae1-a738-cb20d5ee4566 · outbound

This paper cites VALTEST: Automated validation of language model generated test cases.arXiv preprint arXiv:2411.08254, 2024.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets VALTEST: Automated validation of language model generated test cases.arXiv preprint arXiv:2411.08254, 2024

Reference 55

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:43:01.743085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:43:00.527312Z digest=sha256:424b1f7f402439b2b1ec2aaaa8ecb84fc14fde00ca06bdced6dc8847af660c9a

Observation 6afcd0e0-4214-4d8a-b992-712d4c9a5c4e · outbound

This paper cites Hallucination to consensus: Multi- agent LLMs for end-to-end JUnit test generation.ACM Transactions on Software Engineering and Methodology, 2026.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Hallucination to consensus: Multi- agent LLMs for end-to-end JUnit test generation.ACM Transactions on Software Engineering and Methodology, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.589660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.589660Z digest=sha256:19e8667c045f15b1c3a677ee0d563458f22ca55f192a8bc5ae677de1afa29627

Observation 02d73510-88a2-49a8-8ad6-887c67f30d77 · outbound

This paper cites Do LLMs generate test oracles that capture the actual or the expected program behaviour?.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Do LLMs generate test oracles that capture the actual or the expected program behaviour?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.711348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.711348Z digest=sha256:770536059399ea515273be6e53b3f843751f98cc5004c309fea1954e33fb1826

Observation b6b194e9-ca16-465e-a061-06b656f02174 · outbound

This paper cites PAL: Program-aided Language Models.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets PAL: Program-aided Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.801354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.801354Z digest=sha256:36d969fbf3802c652c73cd21e8c50fd4e6eb8d2ffb92bfaba8814bdc29c88f27

Observation 246a7c5b-3fb0-4a73-a6a8-40f2cfbd12ef · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.878591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.878591Z digest=sha256:5f9591634148bbb3b5e2b6016e14307a36508286cd176217adc67bcbba727968

Observation fcb0a385-ed47-474e-88e8-9680e753f793 · outbound

This paper cites TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.923422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.923422Z digest=sha256:077788fb5b82f88c99813a74b363fb6f81b53cbe35995e87f48afdd93f2cd097

Observation 2a2c8d15-daa5-437a-bf31-b20394ddef9d · outbound

This paper cites Sequential enumeration in large language models.arXiv preprint arXiv:2512.04727, 2025.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Sequential enumeration in large language models.arXiv preprint arXiv:2512.04727, 2025

Reference 61

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:43:01.425149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T00:43:01.079571Z digest=sha256:2286224d32fbc0043c32b6823f1328315471840284067c3fad474121ead56a40

Observation 59fdd9e4-6fcb-4ed3-a5c9-5609e1b3485a · outbound

This paper cites |S ∗|= 10.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets |S ∗|= 10

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:01.149558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:01.149558Z digest=sha256:61cad0ab09b6acd1f01c412916a6b14d96a085321d9b000041f704235a07548b

Pith citing papers

No inbound Pith citation observations are available.