Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:43:01.149558Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.01000.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:43:01.149558Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56765134-c1f1-45e3-8ad6-3b61f52374e3 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Concrete Problems in AI Safety
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b5ba70-1fdd-497c-aafb-e2b2127b7b1f · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b80d79f-183c-4525-94a0-aabd8181c5d5 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c967abae-1ab3-4368-b4b6-5466d675666d · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6850ce16-9810-4849-b0c8-5a01b6b1b022 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets CodeT: Code generation with generated tests
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a6f23190-75dd-4ea2-a758-ac7ab7893f39 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80aa02aa-e708-4211-a92a-734accb60b35 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Jimenez, John Yang, Kevin Liu, and Aleksander Madry
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f09a48a7-2071-4fff-9943-f16d5e6aaff7 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ee833d-e729-49cc-ac06-0519fce26530 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005b98e3-25d7-4240-a9d3-51f1bbc8ddb5 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Rabe, Talia Ringer, and Yuriy Brun
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e545781d-73f3-4ba4-ad35-7054ea7994a2 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Scaling laws for reward model overoptimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 801c6585-745f-478f-b774-0563d4008229 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets A Survey on LLM-as-a-Judge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c120526-1391-40cd-8620-153f1f432da0 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbdc489b-57d3-41ac-8f0c-cea0e5b2814b · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c547609-724c-468b-ac3d-2d5426bc755a · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Large language models cannot self-correct reasoning yet
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a3212086-ea3f-4470-b589-c33e24fd66c1 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Pitfalls of rule- and model-based verifiers: A case study on mathematical reasoning.arXiv preprint arXiv:2505.22203, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04fee6b-eb3d-492a-8552-fb573dc0c5a7 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d63866-d6d2-445d-90d9-125f3954079a · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a89ff8da-017e-4a43-8066-a445deb8ca2d · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Language Models (Mostly) Know What They Know
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b5bf56-2455-40df-8c27-4ac6a37a2744 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d8bc4f-2ab0-47ad-996f-6442d3c29fbb · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Let’sverifystepbystep
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1041a728-0780-4d1b-9b1e-9fc7e41be5ef · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f45f340b-6e35-426b-a2c7-6c37c87c9d3c · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Rethinking Verification for LLM Code Generation: From Generation to Testing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 261b4d69-af70-4247-8968-b61cb2c8d588 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-refine: Iterative refinement with self-feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5c55bc9d-56f5-495b-8c7a-f85f33a27070 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d4dedd60-c8fd-44a7-ae32-aa71ce43d2a6 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets What can we learn from collective human opinions on natural language inference data? InProceedings of EMNLP, 2020
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ae0cc0b-53c8-4a2e-8a68-5bcdab945dd1 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets The effects of reward misspecification: Mapping and mitigating misaligned models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7d0c9a5-fc65-4dde-ac36-f927e43501aa · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Qwen2.5 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5f5215-d663-4eaa-8ade-5d68212353fb · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Before the Model Learns the Bug:Fuzzing RLVR Verifiers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12ec88ad-73c2-45b3-82ff-215cc94ad8ef · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets An empirical evaluation of using large language models for automated unit test generation.IEEE Transactions on Software Engineering, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4d1e0d3c-c731-42f6-81d8-0aa460389574 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59fe8eca-91b5-4b3b-b61f-47f671f12e7b · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a151dce-1fbe-4ea6-af68-2b1378e91128 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f457223-c64b-4f60-b8b0-d9575cbc3583 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683e9c94-b538-447e-bfd9-f61edfaee109 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 32f3b319-5db9-4391-9e00-d3e90967dfd7 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Order matters: Sequence to sequence for sets
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 290ed08f-de5f-440a-88ac-1ec852d72c44 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-consistency improves chain of thought reasoning in language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 99421d6f-cc4f-4d6f-ac40-2c5440992a10 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Con- sistency of a recurrent language model with respect to incomplete decoding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2cc989b6-bc3c-4596-a067-a6cc37a92e28 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Large language models are better reasoners with self-verification
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a972c1a-5c23-46e9-92af-72b9315f088c · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets what it can create, it may not understand
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2df6edbf-9234-47ab-abbb-249b9208a7e8 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Efficient Guided Generation for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c82dbb-8a4e-48cb-950a-43421840574c · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Jiang, Wenda Li, et al
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 18f93d2b-342f-4e6e-9566-7a552c9195ed · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0e153ff8-2c8e-4109-a9a0-2c5f7b589980 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c690072c-e9ad-48de-9f5c-32e63bac7c6b · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Judging LLM-as-a-judge with MT-bench and chatbot arena
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 047a85fc-af93-4817-91e6-e91dc3018dd5 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets RubricBench: Aligning model-generated rubrics with human standards
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7c10b3-f874-4fe8-8bb0-b59dfb1efac2 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-Rewarding Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab446688-d7fe-443e-ab53-2f9e60002670 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12947eb9-15c3-41cd-ac6d-73160fddd018 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets LLM Critics Help Catch LLM Bugs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796a9f74-3c3c-46c4-898b-aedb632ee5ab · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets SWT-Bench: Test- ing and validating real-world bug-fixes with code agents
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9a5417b1-96e7-432a-b22a-eeb1cf7271d2 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7e78a78f-6223-4c63-9ea3-3b09e37e95c8 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Tomayto, tomahto: Beyond token-level answer equivalence for question answering evaluation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5cd6e2e8-5672-4516-a13f-9301b2efc1fc · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TOGLL: Correct and Strong Test Oracle Generation with LLMs
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec88bab9-d4f6-4ae1-a738-cb20d5ee4566 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets VALTEST: Automated validation of language model generated test cases.arXiv preprint arXiv:2411.08254, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6afcd0e0-4214-4d8a-b992-712d4c9a5c4e · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Hallucination to consensus: Multi- agent LLMs for end-to-end JUnit test generation.ACM Transactions on Software Engineering and Methodology, 2026
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d73510-88a2-49a8-8ad6-887c67f30d77 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Do LLMs generate test oracles that capture the actual or the expected program behaviour?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b194e9-ca16-465e-a061-06b656f02174 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets PAL: Program-aided Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246a7c5b-3fb0-4a73-a6a8-40f2cfbd12ef · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb0a385-ed47-474e-88e8-9680e753f793 · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a2c8d15-daa5-437a-bf31-b20394ddef9d · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Sequential enumeration in large language models.arXiv preprint arXiv:2512.04727, 2025
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 59fdd9e4-6fcb-4ed3-a5c9-5609e1b3485a · outbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets |S ∗|= 10
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.