Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:59.087204Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.24098.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:59.087204Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:58:27.773874Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:00:08.065119Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d8a05b95-d4d3-4b7c-b4df-3ee993ddfcb5 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding 1 1 0\n1000000000\n1000000000
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e799839a-3f33-456d-93ae-86b8d820e280 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding 123 * The Python code block under each field should be independent
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 061501b8-bf53-43ab-a118-6973da65e000 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e236a05-f216-4964-b56f-286e65daac9d · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding {n} { m}
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2df2e97a-34d1-494c-8141-6e19422bf7fb · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Measuring Coding Challenge Competence With APPS
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e776c6e1-13d8-4fcc-b270-a166d4077ee4 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24efad3f-0eec-4e62-8e7b-b58fb412704c · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c7557f-7a31-4555-a1dc-c81aa76c3ac5 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20512dfd-441a-43d1-8ede-44926968ecd0 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competition-Level Code Generation with AlphaCode
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf60f56-81a0-44af-8315-8fc8bfe65d73 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Scattered Forest Search: Smarter Code Space Exploration with LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d06d85-d232-48cf-8b08-7e612b513762 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe2febf-5736-49ce-ae1f-f9eefc4f3563 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202c5d0d-e4de-49bf-9e6e-a487bf6ee355 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding URL http://dx.doi.org/10.1145/3510454.3516829
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89f2e36-8c6d-44cb-a848-bf6604ace09f · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenAI o1 System Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82330be4-dce1-4de6-8fb7-88075f7db129 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Competitive Programming with Large Reasoning Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa34b9b4-ffad-4aa3-83ca-aa3829bb7f6f · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding COFFE: A Code Efficiency Benchmark for Code Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca57237a-acf1-443b-a1e0-ad4fabb214be · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d92569e-dd0a-41c5-9707-ddcccb60ff5d · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding TESTEVAL: Benchmarking Large Language Models for Test Case Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61da6a1-b863-4077-b298-007329eaa1e2 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b65fe22-9ec2-4d7a-b4b4-46c5125f8fd9 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMO: Less is More for Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938c11e9-cd80-453e-88aa-75b304857657 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0285f0-db1e-4107-888c-a8124f293355 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe0e6ca-7484-47d1-a1e4-74c89c72c79e · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding ALGO: Synthesizing Algorithmic Programs with LLM-Generated Oracle Verifiers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05460c55-7a88-42d1-a9f5-469868613ccd · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74403fe-084f-422b-9b8d-5d4808eebf4f · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0d58dc-98f6-4d04-abb3-7da3faa8eed2 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dcfff285-cc90-4ffd-a491-3123e6f4879e · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding core logic
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d69afd2e-7914-4815-a77d-59fb5b820682 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding 3\n1 10\n2 8\n3 10
Reference 1000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c47beb24-a9e0-4f8b-8614-f0da8e6c0637 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding ISBN 9781450304436
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5115e784-7837-4ed2-986a-be32a29beb91 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding Program Synthesis with Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d72edb9-f7f9-44e4-8f82-d98ba46e6223 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding TACO: Topics in Algorithmic COde generation dataset
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f145b96d-c545-4261-bb99-07a16a9a238b · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51511022-9523-4e30-8b97-5c23997f1d84 · outbound
HardTests: Synthesizing High-Quality Test Cases for LLM Coding OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57da5526-4a26-4dbb-8eed-1f45a9385867 · inbound
Efficiency of turbulence HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb1423bc-89a7-4034-949f-64dd43891641 · inbound
Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02a07b45-6973-4a29-84d8-0db561655f4d · inbound
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f72180a-fb8a-48f2-93ee-2b50f2e2cfc5 · inbound
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6f235e1-2126-485a-a8bb-23a8c8e84267 · inbound
CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99f18f25-3205-49c3-8e8a-86caaa7b88ff · inbound
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b618dcfd-225a-4794-85de-af1428477042 · inbound
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0aa7b6c7-458a-4bf8-a924-98bbb9ae9c7b · inbound
Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.