Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2412.21199.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:54:20.105337Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T14:33:31.671093Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c1dc0812-c573-420c-a324-23f2aca5ef02 · inbound
Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e3b6f0-bf66-4c79-a1bc-9e4cc155d542 · inbound
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · inbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9a34e2-15bc-46cd-893c-47c92f4f07ed · inbound
LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d81c616b-3c83-48c7-8893-5aa8fa3fb860 · inbound
Agentic Frameworks for Reasoning Tasks: An Empirical Study HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae7d3cc9-89ac-4a6d-bc98-253c89ec33fc · inbound
Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4524c26e-f802-4aa7-888b-6e408c19af54 · inbound
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60f2ccbc-e987-460d-a2fc-12697d43e592 · inbound
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d49d5d0-ece7-4be7-bb7e-e28556ea6543 · inbound
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e3c6695-b408-4083-a1cc-6c9803c17d38 · inbound
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.