Pith. sign in

Paper Citation Record · LEDGER

HardTests: Synthesizing High-Quality Test Cases for LLM Coding

As of 15 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.24098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24098 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:59.087204Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:58:27.773874Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.065119Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8a05b95-d4d3-4b7c-b4df-3ee993ddfcb5 · outbound

This paper cites 1 1 0\n1000000000\n1000000000.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 1 1 0\n1000000000\n1000000000

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.263496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:39:59.087204Z digest=sha256:e4ed9b05d8264ac3cc39132a0ef5f8fe2109a163d77ca57d7e741b6868e4f946

Observation e799839a-3f33-456d-93ae-86b8d820e280 · outbound

This paper cites 123 * The Python code block under each field should be independent.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 123 * The Python code block under each field should be independent

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.825078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:39:58.827514Z digest=sha256:d75a902c8656027a874c035563786d1dee849090f4810a3e9578af07bb5e9bd0

Observation 061501b8-bf53-43ab-a118-6973da65e000 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.616900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.616900Z digest=sha256:ee8c2275954424f8730d0a88d99a642ba66abf43a52b742c2394dd71a7fba656

Observation 1e236a05-f216-4964-b56f-286e65daac9d · outbound

This paper cites {n} { m}.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding {n} { m}

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.424892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:39:59.017463Z digest=sha256:92bce953e576b5807e1e811e115ecb883ff2669c45b062af9a4728b00cbbd4da

Observation 2df2e97a-34d1-494c-8141-6e19422bf7fb · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Measuring Coding Challenge Competence With APPS

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.896221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.896221Z digest=sha256:d1f1491a0f19bc2d596b17270642ba0b9e289de99a00753e0253bd06ed5fa456

Observation e776c6e1-13d8-4fcc-b270-a166d4077ee4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.068560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.068560Z digest=sha256:088409f056ad30c4044d78f0ecacf87bac11389e426deef57c0ff921e776d6e0

Observation 24efad3f-0eec-4e62-8e7b-b58fb412704c · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.138659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.138659Z digest=sha256:9a1da99c792cdaea43037d0bc9f50bd89c577d2cabd6c62c6c803809e0facc3c

Observation 82c7557f-7a31-4555-a1dc-c81aa76c3ac5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.209649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.209649Z digest=sha256:b8f277fca22c552f8218d62b7fec4a518dd4d6bac337b8de64be5dc594853414

Observation 20512dfd-441a-43d1-8ede-44926968ecd0 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competition-Level Code Generation with AlphaCode

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.557357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.557357Z digest=sha256:6cb5570738d535a1722b76cc2977a8fb94f12eb31c49a5059ce9beba22872623

Observation 6bf60f56-81a0-44af-8315-8fc8bfe65d73 · outbound

This paper cites Scattered Forest Search: Smarter Code Space Exploration with LLMs.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Scattered Forest Search: Smarter Code Space Exploration with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.669431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.669431Z digest=sha256:312e74840e13c243a4d99d6a28b78efec4b71db4ab16c6a4b626132d122c38ad

Observation 87d06d85-d232-48cf-8b08-7e612b513762 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.781297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.781297Z digest=sha256:7a24f84f62232af0491b7095e85040affa83acb57f5c8f99f25452c186650935

Observation 4fe2febf-5736-49ce-ae1f-f9eefc4f3563 · outbound

This paper cites rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.905741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.905741Z digest=sha256:26defd184c480794d04345da24350e2f540df4c28b2a0a60648f41bbd8944072

Observation 202c5d0d-e4de-49bf-9e6e-a487bf6ee355 · outbound

This paper cites URL http://dx.doi.org/10.1145/3510454.3516829.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding URL http://dx.doi.org/10.1145/3510454.3516829

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.037173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.037173Z digest=sha256:839f4b691e5d3eac00b9074a8e792cb1e3a78251b948675a027102e36cde0c2d

Observation d89f2e36-8c6d-44cb-a848-bf6604ace09f · outbound

This paper cites OpenAI o1 System Card.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.197862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.197862Z digest=sha256:32b4ba36dd0331edf5264ec76216c1e8bd78bdeaf93e4689638803092ef54b0d

Observation 82330be4-dce1-4de6-8fb7-88075f7db129 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competitive Programming with Large Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.365862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.365862Z digest=sha256:a6d71dd67efed8a5647f5f48340f06a58fe79e332d734a1a184a768e31facaaa

Observation aa34b9b4-ffad-4aa3-83ca-aa3829bb7f6f · outbound

This paper cites COFFE: A Code Efficiency Benchmark for Code Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding COFFE: A Code Efficiency Benchmark for Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.465250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.465250Z digest=sha256:ff4293c5216a14502c94ffccb99f855832d2507eecf551dd3492a15d731a1f68

Observation ca57237a-acf1-443b-a1e0-ad4fabb214be · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.560048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.560048Z digest=sha256:9a64bab7193667a0b88d291d289b8afc1601f06e397f7f690ec4528f3fc115fc

Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.636759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.636759Z digest=sha256:09f51d8a21f4d7baffaa55db0b355f86e89342d1c685bf66f49bd2e60f79139a

Observation 9d92569e-dd0a-41c5-9707-ddcccb60ff5d · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.721681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.721681Z digest=sha256:21148fb8431559bc33571a4959050b5136b74b14c6f7917027f633d25a5e9b01

Observation c61da6a1-b863-4077-b298-007329eaa1e2 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.824234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.824234Z digest=sha256:e1b11b120109441f2e671a673c2638b05a87c4ddcc138c08d0f26c1eb66f49e4

Observation 8b65fe22-9ec2-4d7a-b4b4-46c5125f8fd9 · outbound

This paper cites LIMO: Less is More for Reasoning.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMO: Less is More for Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.934398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.934398Z digest=sha256:383bd076f71ede950303f604611c357b548282f346371f623f7c3bc073dbba1e

Observation 938c11e9-cd80-453e-88aa-75b304857657 · outbound

This paper cites No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.040594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.040594Z digest=sha256:6d73c32eed00206d7d921033d49a518a7ccaeb558bbfe5a1d2a427bd672890b5

Observation ae0285f0-db1e-4107-888c-a8124f293355 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.116873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.116873Z digest=sha256:cd54236c6bfb0cc4cc90a40b12216dda8db93f86bccca0fccb457815d3797e76

Observation 7fe0e6ca-7484-47d1-a1e4-74c89c72c79e · outbound

This paper cites ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.211974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.211974Z digest=sha256:6e0da4a2d844dea818a1460aaeb9fed11a8bb1e0dac93d02a79da06a8ab58d99

Observation 05460c55-7a88-42d1-a9f5-469868613ccd · outbound

This paper cites TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.370843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.370843Z digest=sha256:70e2e0dbc1ca302157d51643d3cebf8b241d7b8504b1ee3fff556fe66519572d

Observation e74403fe-084f-422b-9b8d-5d4808eebf4f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.472965Z digest=sha256:7a7b88832ffa9b2b0bbf43f700b680e62a6a133e67c6e9a33e682c7d6ad1e7ea

Observation 5a0d58dc-98f6-4d04-abb3-7da3faa8eed2 · outbound

This paper cites an unresolved cited work.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:01.099920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:39:58.578150Z digest=sha256:924048542a43152dc6de1372300c69159e79eed80db71c3a03343ef55774b9d5

Observation dcfff285-cc90-4ffd-a491-3123e6f4879e · outbound

This paper cites core logic.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding core logic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.965186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:39:58.714694Z digest=sha256:d916288339d585e5ca62700892d20d408575c7cc5871d0f275ab5702cdf283a4

Observation d69afd2e-7914-4815-a77d-59fb5b820682 · outbound

This paper cites 3\n1 10\n2 8\n3 10.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding 3\n1 10\n2 8\n3 10

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:00.611877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:39:58.905686Z digest=sha256:655e00c4df578ba625bfa60d8499b1695a54d627724bf033fcfa3ddeb89e3707

Observation c47beb24-a9e0-4f8b-8614-f0da8e6c0637 · outbound

This paper cites ISBN 9781450304436.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding ISBN 9781450304436

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.691285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.691285Z digest=sha256:ab8e6ffe3162b2cc5019a8b84514efd0178715ad15f01143bd2ecd817b9f6e50

Observation 5115e784-7837-4ed2-986a-be32a29beb91 · outbound

This paper cites Program Synthesis with Large Language Models.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Program Synthesis with Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.538922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.538922Z digest=sha256:784af7355fd5bd2d28e8eedcc2174e7518cb0de631d0194a3d5c311669c24418

Observation 5d72edb9-f7f9-44e4-8f82-d98ba46e6223 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TACO: Topics in Algorithmic COde generation dataset

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.321057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.321057Z digest=sha256:68d7e8266cd4327307866dbb4ebd64d4376ba77e16a2b6ac90e44740877fc947

Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · outbound

This paper cites LIMR: Less is More for RL Scaling.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.430228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.430228Z digest=sha256:e278677ba8efc69aadc5dbc519ec55d5f2bec75a651b352d688d69d7618e004e

Observation f145b96d-c545-4261-bb99-07a16a9a238b · outbound

This paper cites TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.474805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.474805Z digest=sha256:69d104c1df81411859a659f2c3654a1f71989aecfa1ffd53b0b1197a881764b0

Observation 51511022-9523-4e30-8b97-5c23997f1d84 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:55.391872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:55.391872Z digest=sha256:ebaffb3d9ce6abd0816eaaee5f67a612a7d82caa69999896749fcd4bbcb12dfa

Pith citing papers

Observation 57da5526-4a26-4dbb-8eed-1f45a9385867 · inbound

Efficiency of turbulence cites this paper.

Efficiency of turbulence HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:58:27.773874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:58:27.773874Z digest=sha256:6a5d46a697179b97aeb953a6c657ff7991027b9bfdd6815fb42c91533919ac75

Observation bb1423bc-89a7-4034-949f-64dd43891641 · inbound

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems cites this paper.

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:05.487820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T23:27:54.704794Z digest=sha256:f5c35a976727447f6698180f60dfd4d1af4e800d97f8e069f94513488bf739c2

Observation 02a07b45-6973-4a29-84d8-0db561655f4d · inbound

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation cites this paper.

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:32.448442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T01:21:42.562823Z digest=sha256:d55591f58a5eb665cc2e552d0ae7da85155c992b08ca96904bbca86c9ee87689

Observation 7f72180a-fb8a-48f2-93ee-2b50f2e2cfc5 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.053765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:ce170ce5830ea7bcaaddf485d0696d494b2a24e61aa6f82499e0e5b60fc520e4

Observation c6f235e1-2126-485a-a8bb-23a8c8e84267 · inbound

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming cites this paper.

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.201084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T13:19:22.541146Z digest=sha256:d01aa4253f8a8d1beefb2decb4fddbceec7b4e733a080e9799a4b606c1308919

Observation 99f18f25-3205-49c3-8e8a-86caaa7b88ff · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.067026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T20:55:15.784610Z digest=sha256:10c40a1da9cd14c5577e03c6795d4e0c40842e1158ad81c12ed87b61e91d3ba5

Observation b618dcfd-225a-4794-85de-af1428477042 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.596811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:29:21.598397Z digest=sha256:0791841e29a1daf161a33d18925be046ae65b1c2bed7ca18820aa7dd20e7afd5

Observation 0aa7b6c7-458a-4bf8-a924-98bbb9ae9c7b · inbound

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch cites this paper.

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch HardTests: Synthesizing High-Quality Test Cases for LLM Coding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:23:16.109674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:23:16.109674Z digest=sha256:1c7de259669abdabeee77fd698c08bd9827081e42a3d116e690418562d070f50