Pith. sign in

Paper Citation Record · LEDGER

HardTests: Synthesizing High-Quality Test Cases for LLM Coding

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.24098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24098 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:59.087204Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:58:27.773874Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.065119Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8a05b95-d4d3-4b7c-b4df-3ee993ddfcb5 · outbound

This paper cites 1 1 0\n1000000000\n1000000000.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 1 1 0\n1000000000\n1000000000

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.263496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:59.087204Z digest=sha256:5892e23dbf3a564b5ed7343152ae38619f3557a90db1f7f71fb12744a1de92db

Observation e799839a-3f33-456d-93ae-86b8d820e280 · outbound

This paper cites 123 * The Python code block under each field should be independent.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 123 * The Python code block under each field should be independent

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.825078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.827514Z digest=sha256:8d0d3241c2c0b306eed7d30c33b8130a37e662a7c4f6e418ca37f96ae77c5d07

Observation 061501b8-bf53-43ab-a118-6973da65e000 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.616900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.616900Z digest=sha256:15ba67bbba251795c1f828511c5e9da5661fcf71a686234fc40389baf3b5c461

Observation 1e236a05-f216-4964-b56f-286e65daac9d · outbound

This paper cites {n} { m}.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding {n} { m}

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.424892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:59.017463Z digest=sha256:86fa3b8f495f9f1c031eef7a5b9a9dd39c3b70b0faa241f33402a6f57574df97

Observation 2df2e97a-34d1-494c-8141-6e19422bf7fb · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Measuring Coding Challenge Competence With APPS

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.896221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.896221Z digest=sha256:3cbdbb8867f165710d58b3a10a418b5b2e1695e0108995a65ed0332ce110f5f5

Observation e776c6e1-13d8-4fcc-b270-a166d4077ee4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.068560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.068560Z digest=sha256:36bedc0ee5944de3f125be79f826e96019b2c1e49e5e37f19606e026a64dbb1b

Observation 24efad3f-0eec-4e62-8e7b-b58fb412704c · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.138659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.138659Z digest=sha256:37851b14c286b4d62e39675fd30ed57816e0799f1a16b3aca94660791eeac3de

Observation 82c7557f-7a31-4555-a1dc-c81aa76c3ac5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.209649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.209649Z digest=sha256:ee19dfb414e39efc39072f9b584c0a3dc12056b6d43a6e5ec0e3bf9f2bd906a0

Observation 20512dfd-441a-43d1-8ede-44926968ecd0 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competition-Level Code Generation with AlphaCode

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.557357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.557357Z digest=sha256:08623ba5cd4f9e253e439acb685ca55cb9df47226d993088af2799752eeb7546

Observation 6bf60f56-81a0-44af-8315-8fc8bfe65d73 · outbound

This paper cites Scattered Forest Search: Smarter Code Space Exploration with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Scattered Forest Search: Smarter Code Space Exploration with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.669431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.669431Z digest=sha256:4b5e477e78fbba924823c248c7cf899e58d2f9176f777b431914950b429e3cc1

Observation 87d06d85-d232-48cf-8b08-7e612b513762 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.781297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.781297Z digest=sha256:5c7731002b4673ae2a65536bb8c863334422bd420624e848f679fd5542cfaf61

Observation 4fe2febf-5736-49ce-ae1f-f9eefc4f3563 · outbound

This paper cites rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.905741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.905741Z digest=sha256:3d62bc1aafe114fc1955573ed11311721cf85f999bc44f6548e56c47b33477f0

Observation 202c5d0d-e4de-49bf-9e6e-a487bf6ee355 · outbound

This paper cites URL http://dx.doi.org/10.1145/3510454.3516829.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding URL http://dx.doi.org/10.1145/3510454.3516829

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.037173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.037173Z digest=sha256:92f638f221f2a26b8181338da612d5c76daeb29fad64dcf1fd2bf93df72c6550

Observation d89f2e36-8c6d-44cb-a848-bf6604ace09f · outbound

This paper cites OpenAI o1 System Card.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.197862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.197862Z digest=sha256:82a88f45d84a47003103f0b50c744612e0bde5fd8deb384b89fd15429314174f

Observation 82330be4-dce1-4de6-8fb7-88075f7db129 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competitive Programming with Large Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.365862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.365862Z digest=sha256:ab51057d68edcbf957a31b4b36e789b76ebee05750d370a662f7f6b83260ee51

Observation aa34b9b4-ffad-4aa3-83ca-aa3829bb7f6f · outbound

This paper cites COFFE: A Code Efficiency Benchmark for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding COFFE: A Code Efficiency Benchmark for Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.465250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.465250Z digest=sha256:1e268c403e8ff798f10419f4142fd5b176c967943346255a34a505fbc00b140d

Observation ca57237a-acf1-443b-a1e0-ad4fabb214be · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.560048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.560048Z digest=sha256:29292d49c31e7d7fe5227e6f1bdcd0538680f162005ac6e593faf514b46e9087

Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.636759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.636759Z digest=sha256:e3ebc2cbef545305f31f5b99b75ffa0cbba3eb411aa81c1b4c7ecaf993d86b95

Observation 9d92569e-dd0a-41c5-9707-ddcccb60ff5d · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.721681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.721681Z digest=sha256:f763165bb0b5dcdfd5845075b27d416475ea5798fb2c99d314a380a96340b1f8

Observation c61da6a1-b863-4077-b298-007329eaa1e2 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.824234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.824234Z digest=sha256:4e982c35fa066b21d8c1df2e1a683dcf3a519255b49ae62643c9eb33f6016e4f

Observation 8b65fe22-9ec2-4d7a-b4b4-46c5125f8fd9 · outbound

This paper cites LIMO: Less is More for Reasoning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMO: Less is More for Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.934398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.934398Z digest=sha256:88d98d6d1140be5c9cd953a0b6abf7e557bac3ad9123ac07f42334f7abb06ca2

Observation 938c11e9-cd80-453e-88aa-75b304857657 · outbound

This paper cites No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.040594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.040594Z digest=sha256:1354ee4fc31f0bdd34e3307e97e75c7e9dfa82316401466a0016debe13c2adb5

Observation ae0285f0-db1e-4107-888c-a8124f293355 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.116873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.116873Z digest=sha256:73030b50e7ccd6934f70f86809153e543923c1494ec38b2266edecfd739fa952

Observation 7fe0e6ca-7484-47d1-a1e4-74c89c72c79e · outbound

This paper cites ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.211974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.211974Z digest=sha256:607a85cb0da0094aff67b17277500422703bede71f650f6bcddfebab6078d165

Observation 05460c55-7a88-42d1-a9f5-469868613ccd · outbound

This paper cites TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.370843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.370843Z digest=sha256:94db769e0a2583d50e23f660b8d2da99448092431531fbc4c18a1a6b1f8e150e

Observation e74403fe-084f-422b-9b8d-5d4808eebf4f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.472965Z digest=sha256:3ec20b50157d998dfe135be8f338c9d39a953c990da771345bb842180f6c97e9

Observation 5a0d58dc-98f6-4d04-abb3-7da3faa8eed2 · outbound

This paper cites an unresolved cited work.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:01.099920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.578150Z digest=sha256:6a2bb8d8f99249678d7763c93e7d8424c4cca0b166b68dfe84cfdd325ad224e0

Observation dcfff285-cc90-4ffd-a491-3123e6f4879e · outbound

This paper cites core logic.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding core logic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.965186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.714694Z digest=sha256:043a6ecc987d2d768efa60a0d2005cb293ab73d2d6019742b3b3ade321936f12

Observation d69afd2e-7914-4815-a77d-59fb5b820682 · outbound

This paper cites 3\n1 10\n2 8\n3 10.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 3\n1 10\n2 8\n3 10

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.611877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:39:58.905686Z digest=sha256:3eb39144f25eb74589450804b2a04be48777dc94840cb72057376e577ef60789

Observation c47beb24-a9e0-4f8b-8614-f0da8e6c0637 · outbound

This paper cites ISBN 9781450304436.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ISBN 9781450304436

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.691285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.691285Z digest=sha256:12a46e1ca6d8065db632ba9d081089ea651b3e2f6eb3d5d50f264a96454e02b9

Observation 5115e784-7837-4ed2-986a-be32a29beb91 · outbound

This paper cites Program Synthesis with Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Program Synthesis with Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.538922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.538922Z digest=sha256:d54e2dd0431e13eb6fcdff6a74b706ef3c35865a483dbba1180493edef98531b

Observation 5d72edb9-f7f9-44e4-8f82-d98ba46e6223 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TACO: Topics in Algorithmic COde generation dataset

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.321057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.321057Z digest=sha256:0958100ffecd51c61cd40c477b274b2c992c33fbd5dad8f3e0c557828c3f849e

Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · outbound

This paper cites LIMR: Less is More for RL Scaling.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.430228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.430228Z digest=sha256:070bbd85a96f6f7e320e3379d4547437ca1ef82c193d723feceedaf3eb784c65

Observation f145b96d-c545-4261-bb99-07a16a9a238b · outbound

This paper cites TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.474805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.474805Z digest=sha256:8430d4e993abddfaa571471526e7c74b201bad25649f702316421b0cedf99f7f

Observation 51511022-9523-4e30-8b97-5c23997f1d84 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.391872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.391872Z digest=sha256:305554d19708d881957b1792e38d482fb7079f30dc0162c0c8daa43ba1d69a0e

Pith citing papers

Observation 57da5526-4a26-4dbb-8eed-1f45a9385867 · inbound

Efficiency of turbulence cites this paper.

Efficiency of turbulence HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:58:27.773874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:58:27.773874Z digest=sha256:35e71615759a1b535735344991bcf35f996b2b700bf53bc659b3eb8ca4b96504

Observation bb1423bc-89a7-4034-949f-64dd43891641 · inbound

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems cites this paper.

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:05.487820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T23:27:54.704794Z digest=sha256:134192ce93886caf6dab43c2dede60f9401bd961f679ebeb4486bb93bff1b376

Observation 02a07b45-6973-4a29-84d8-0db561655f4d · inbound

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation cites this paper.

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:32.448442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:21:42.562823Z digest=sha256:99959f3020ecf9b3b5d91fd99d9b1c958bac513182bbbfe2dad1f0e0edff7246

Observation 7f72180a-fb8a-48f2-93ee-2b50f2e2cfc5 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.053765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:e58bcee083b678f59b9e67fb6ac170806948bc639eb33a9af8880e039d4fdd3f

Observation c6f235e1-2126-485a-a8bb-23a8c8e84267 · inbound

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming cites this paper.

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.201084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:19:22.541146Z digest=sha256:d7d23c1efeeee98dd41c8aad2be1ac134847900b3afa09d922bd4e74af602f7b

Observation 99f18f25-3205-49c3-8e8a-86caaa7b88ff · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.067026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:55:15.784610Z digest=sha256:1ba49fbd303bcd45c48d86767ef71abd52b2b3c2961881c2c87e0d0ec1f9e7f0

Observation b618dcfd-225a-4794-85de-af1428477042 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.596811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:29:21.598397Z digest=sha256:64e353786b3220a193852852bf2f0cac8291fead38afdb1bf07b84d991366ba9

Observation 0aa7b6c7-458a-4bf8-a924-98bbb9ae9c7b · inbound

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch cites this paper.

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:23:16.109674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:23:16.109674Z digest=sha256:6aa17819531456de2082df234159f6b27eee8a51b3e5cedfc3fc4fbfe2dfb541