Pith. sign in

Paper Citation Record · LEDGER

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2412.21199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21199 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:54:20.105337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:33:31.671093Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1dc0812-c573-420c-a324-23f2aca5ef02 · inbound

Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models cites this paper.

Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:20.105337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:20.105337Z digest=sha256:8b4272e477ba950bfc9c0500e688122726d5b2e024012a587b2f3ffa9b80e501

Observation 87e3b6f0-bf66-4c79-a1bc-9e4cc155d542 · inbound

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure cites this paper.

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:57.050750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:57.050750Z digest=sha256:cc57a6086c18a70a24419cdd277df5b3927a7fc75d448a95b12256b3799f917d

Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · inbound

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks cites this paper.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.114501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.114501Z digest=sha256:ffeb4e0284e4916307b8c144c792521622011141195a7ba2bc55d8969acb1ae3

Observation ce9a34e2-15bc-46cd-893c-47c92f4f07ed · inbound

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review cites this paper.

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.775143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:52:49.324500Z digest=sha256:edf696336552837100b203f2bd72641620528e9a8850087113f9588160a119ca

Observation d81c616b-3c83-48c7-8893-5aa8fa3fb860 · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:08:26.182245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:f01b819d117bae9bd5bace3af9fb97a47826cb508cf63fc2667ad3f2c68fae00

Observation ae7d3cc9-89ac-4a6d-bc98-253c89ec33fc · inbound

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions cites this paper.

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.331881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:49:57.957009Z digest=sha256:476bc3b606a1e6838fa337a2a4321c4a29162924babb9b1fc4e4409a1e12e068

Observation 4524c26e-f802-4aa7-888b-6e408c19af54 · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:57.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:5b413c274b5076ce8ea6e60771c0a38d02ba61056f86a8f3ff9c5fc4440d3fab

Observation 60f2ccbc-e987-460d-a2fc-12697d43e592 · inbound

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback cites this paper.

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:23:10.609159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T09:22:06.285118Z digest=sha256:b91b355ed6a0c04753e6d22eed7caca262216ea157d6933e9d632d38918bdcec

Observation 7d49d5d0-ece7-4be7-bb7e-e28556ea6543 · inbound

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models cites this paper.

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.672803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T06:25:56.404246Z digest=sha256:afb54e7e061d8aa00c27ced68adafb5e65ec0803d849b64ed83de7f723fdf5ec

Observation 9e3c6695-b408-4083-a1cc-6c9803c17d38 · inbound

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy cites this paper.

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T04:43:45.592808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:43:45.592808Z digest=sha256:085b425f6339c3aacc48a8a08cd12e8af05f0ee121bfbc78b5ec37c21bce2f6c