Pith. sign in

Paper Citation Record · LEDGER

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 11 inbound Pith citation observations for arXiv:2412.21199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21199 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:37.864289Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:46:23.589246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:33:31.671093Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d876dc4f-c33c-4ac7-accd-52835c08fa0a · outbound

This paper cites online" 'onlinestring :=.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.661344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.661344Z digest=sha256:539dac7b3918e8cf0f677a887447993ccb0748e8cf775fc35badc8b2d5a2d66d

Observation 0d88c956-7f5e-47f4-9ad4-f0a8c7f1f6ec · outbound

This paper cites write newline.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.666987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.666987Z digest=sha256:2f9926dcc29079f0240647ab82244929a371fb65abea94df59aa971bfe16a49f

Observation 080a5760-a8d3-474a-b9cc-8212ebbedc5f · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.472932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.673044Z digest=sha256:a6dc9f759b1a82bf4e961bbc26b966fe0592eadcef05111c6f330fdf349a75ae

Observation f8053017-7e59-4e98-8250-3d4326e88a78 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.678098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.678098Z digest=sha256:47caa8862676262cba965b3d8dd70eb07899066f700b5f5d4111f7cce377f4c3

Observation 033749d6-46d4-4338-83ba-565e6f3dc325 · outbound

This paper cites Multi-lingual Evaluation of Code Generation Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.683347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.683347Z digest=sha256:bed97b8c635e5f7f1d7f5b7f31aa82aada92bd916d9f92d10b079d699f5885e0

Observation 081d3982-20cd-4ed4-b84d-6b6827ac32a0 · outbound

This paper cites Program Synthesis with Large Language Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.688411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.688411Z digest=sha256:49893391cea6fe9e6c1377014642cad39543be02af08569755d6ff51681f260f

Observation c6c353f6-ca3b-4472-8e11-8c009960e6c4 · outbound

This paper cites A parallel corpus of Python functions and documentation strings for automated code documentation and code generation.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation A parallel corpus of Python functions and documentation strings for automated code documentation and code generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.692932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.692932Z digest=sha256:19ad02d86aa4a1b9f7b8eb9c11e5db7860f89186aedf35bcaeb7e56824b51cfe

Observation 7a4491f1-207e-4155-8203-e5941a53d6c1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.698545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.698545Z digest=sha256:bee26914dd03db0328c68c4aa56bba137ff21dafbfab21e4d23eb00ed7413ee2

Observation 2c5be030-5340-4c47-aee8-273f73006960 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.704255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.704255Z digest=sha256:f94aca732d94cbff40c1b2783729ce4ee9f3836c507c32169fe99d4730bebb6d

Observation e24e4516-ad1c-4218-9688-cfc5990073b4 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.445866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.709869Z digest=sha256:a19bb57add818f50ca05631f9abdaf3d3244c0183e426cacdd9491dc1168a628

Observation 7e684a6e-11d6-4abd-96f8-514f5772a171 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.715051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.715051Z digest=sha256:46d83d98ad30594bf34c29b36730f827adf6c3a652f25cdc285f785e9ec18dc6

Observation b26248b7-fc10-45f4-97c7-8c4562ad2f08 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.431581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.719215Z digest=sha256:c21dcd730b194dcade7e438e95cfaa876d7e11cd73b94d535186c643882a3e2c

Observation cf1cb287-f0a1-4222-bcd3-33b5898fb3d1 · outbound

This paper cites CoDesc: A Large Code-Description Parallel Dataset.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation CoDesc: A Large Code-Description Parallel Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.723893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.723893Z digest=sha256:667b6bcc7253fd3057367b4e7a23dde5360cefcdd160c2e063ab2449b6962937

Observation 8ce31397-d5bc-4d4f-a982-7e2798b26cba · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.728230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.728230Z digest=sha256:7e282129205194082e7f63d31ce1a12e87e7e59ace1791d4d29abd77d2b50080

Observation 977957ed-a753-4bde-8f22-ffae09f7d481 · outbound

This paper cites Qwen2.5-Coder Technical Report.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Qwen2.5-Coder Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.732820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.732820Z digest=sha256:cd5fefc32b0c4e0e5cc792658e832e820861c6f7522bbdbb34b16cb28d6871a5

Observation d6d5afda-6d22-456c-b57f-bf05192c3346 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.739566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.739566Z digest=sha256:1b024f9777ccd20190b47587b5ded8e093cfe9f73c3e5eddc5d3bdd599f81992

Observation 4546af62-288e-4803-a79f-764d74d0ae2a · outbound

This paper cites Impact of Code Language Models on Automated Program Repair.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Impact of Code Language Models on Automated Program Repair

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.745220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.745220Z digest=sha256:35a82a03e64ce8d34e83b8ca7f25de49f81b662c72c5b09f5769c9db284d50cf

Observation e313ef28-8728-4a80-9525-5fa9c17d9f06 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.749484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.749484Z digest=sha256:495eb031160d8462d9d7bb8645bacec207a556e782f1efab7d4fd3bdb7d781fd

Observation 019ef2cc-36e5-42db-8d4b-fc9d237e9f18 · outbound

This paper cites InferFix: End-to-End Program Repair with LLMs.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation InferFix: End-to-End Program Repair with LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.753760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.753760Z digest=sha256:2af27e6bb046ea0e44dc199e7b7aab1f87dbde0a356c6bddcb8837f2a69c5d99

Observation d898fb23-5cf4-43b7-b2ef-a5be47bca329 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.409903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.758747Z digest=sha256:aecd93a33d4589902b7292422e870aaff5ad4f50488a259f01ab57d2efe1bdd0

Observation d659fbe8-440b-44d8-808f-9e09ed6ac76b · outbound

This paper cites StarCoder: may the source be with you!.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation StarCoder: may the source be with you!

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.762895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.762895Z digest=sha256:c8510a2a83e27faba7127c7abe3cff44cabf9020162d6ca92b7fd3327e0b1147

Observation 8426c663-bc41-40d2-a51c-8d9638c55351 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.394955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.767199Z digest=sha256:c538ad13ed14d45963ef57d1c8179ac325276fd03b116aaf8bdd9323c012f522

Observation 6b7744f3-bc8a-4b86-8175-da019dd24d4b · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.771380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.771380Z digest=sha256:069c0279ce4c3dfdd84f73ee00dc3d71f2ed54c3b7c91041e7ba1c526bff737a

Observation 020fc0de-60e9-4288-9030-e7c11452b449 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.775570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.775570Z digest=sha256:778ee47736bb14fcc244c7c72749d6f6d38bfc07245d96284b731f58b72a83c1

Observation 9546eaaf-9d9c-4d2b-a67e-48fa0fdf7093 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.378797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.780175Z digest=sha256:45e4394cf49a9ba3f03eca9a178d4c9809ab5ad4b7c37de82845e4c48c8eae69

Observation 854ad207-6b18-4fbf-a254-ddb76d19f573 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.784703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.784703Z digest=sha256:243a0fb7803cb3968408b133a55fedc5291e5313f390eefc84a47c19c61f316d

Observation c4169652-c38d-4b82-988e-aa3ec49d70b1 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.789006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.789006Z digest=sha256:26081aee9455032fa3b797357a5551d6ef861e283df1977b8200a3e54638e857

Observation 1748b129-b0bb-4a33-8323-a46721ac04b1 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.353939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.795950Z digest=sha256:4639b398f969468c83313c7a5ccf847142ec1dc27a6dcc3f2d783dde520e75f2

Observation 51a1e376-df4a-47c1-b7e8-86bd9ef400f6 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.329912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.799893Z digest=sha256:6b747023e822779064c6e1da7698d8f7ae3e3cb34e1d5e4053e771db5a5bc1c4

Observation 6fea44e9-578c-491a-8da9-d674ba2718c3 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.804772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.804772Z digest=sha256:aa1029406e65cf271f41fd54fe938319961fd24a47a20f025825b2142494101e

Observation d97a2495-8882-4308-99fc-a8ab6bf6a230 · outbound

This paper cites ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.808644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.808644Z digest=sha256:ab93681a08164bb19236a9eb9348c381b241a16014527ba4d97fb0adc870cb2c

Observation f25c5153-1218-4f6d-a3f0-5ed636b533a0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Code Llama: Open Foundation Models for Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.812756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.812756Z digest=sha256:00d3b629418b386f99e6d6c6f5cfddcec37e9def3353fb44d5c5992299458d27

Observation 66d931f9-994f-445d-85ad-cbe24d7a93f6 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.816642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.816642Z digest=sha256:0ae6455ce87013fc2bca68bacd794455afd90e219d7e83c84b68180ff5716eb5

Observation d639f24a-63f0-48b2-8fca-fa2fc2f95657 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.820346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.820346Z digest=sha256:479dd3eef731b48ef37bcf9eafea00c3fa2bd3c7040d4474219b847b97e14b75

Observation 4492f096-a198-4925-9dad-596d10ecb4d6 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.824782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.824782Z digest=sha256:ded5ceae5e5a9bb5a75befa0c96507133e9ad67ba4b02c569b5265761c5be0d2

Observation 0250df47-431e-4736-9b2c-6263f50e7bdb · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.267828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.828851Z digest=sha256:2cd119346b07a4ff92c61a72444067f104f0769ef969a3c701a5b05c9640902b

Observation 66e6bfa1-eb77-495b-b6a6-c583ae4c3343 · outbound

This paper cites Practical Program Repair in the Era of Large Pre-trained Language Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Practical Program Repair in the Era of Large Pre-trained Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.832818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.832818Z digest=sha256:8838d0492a714b6e7670b3571c5eabe78df71ac0a0ae6e99476ac6b790c7c851

Observation 46bf6f97-003b-4d22-95f6-16f2c6ce59d3 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.836832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.836832Z digest=sha256:af57f6ad30612161c1fb9a8ccb638806325c58cf9d83e1a0377dfd5f8ef81660

Observation 27cb009c-685a-42c3-8b61-3237d64ddafe · outbound

This paper cites Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.840885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.840885Z digest=sha256:8de92db948b5bb17206b43b655da51179b1daac9c3ec949bc7ee10a7d93a5c35

Observation c8abdda8-005d-44d1-a10c-884343b3e9e5 · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.845013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.845013Z digest=sha256:ac4f8442dbceec84dbc34562f605df626f94797ddbaecf58decf51b18c44440e

Observation b15370ef-e38b-4f6e-962c-f751d784b109 · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.849093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.849093Z digest=sha256:4962c1b40d5f74a662987f59f14fe71fc358b480209977f8eb95db9ff08c873b

Observation 5b8c92bb-49a5-4c0b-af41-af0d143da344 · outbound

This paper cites XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.853024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.853024Z digest=sha256:4bd2ca83c0c06d6aac1d2490a7bd2e0758851ccc0d237739a0f9b1de6ea7d42e

Observation fdb3ab70-438b-4c3a-81d7-be9d5ef433e4 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.857588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.857588Z digest=sha256:312ead2274c247e9fd188a358cce833bf907cb8d0e2e2dd68d3d3380a15a4f42

Observation 2bdb5c8e-c31b-4ff2-9365-c2e259feb06c · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.864289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.864289Z digest=sha256:2b7602c4d562f5a5bbd2587e86f863f27a1c507da0cb5d0b27c5fb0a1b2530b9

Pith citing papers

Observation c0188b4d-f77f-400f-b809-c542f336ef2e · inbound

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems cites this paper.

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:23.589246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:23.589246Z digest=sha256:1391f17f711344e46fa7da94f29c98fb48cf9a18843cdec018fe2e16e8fc621c

Observation c1dc0812-c573-420c-a324-23f2aca5ef02 · inbound

Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models cites this paper.

Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:20.105337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:20.105337Z digest=sha256:e5300e47726aaa79b8a32785650d940c7a571a88173e1fdc66c07c3b2a66ff77

Observation 87e3b6f0-bf66-4c79-a1bc-9e4cc155d542 · inbound

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure cites this paper.

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:57.050750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:57.050750Z digest=sha256:cabcaf5dfb385b64a0cfd0c4867ff12f2579006a9752d00b35e1c168db98a51b

Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · inbound

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks cites this paper.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.114501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.114501Z digest=sha256:7c8c27659e80121ec299093c8e2af860e4ce153ad916db972f7d469f724b5aa0

Observation ce9a34e2-15bc-46cd-893c-47c92f4f07ed · inbound

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review cites this paper.

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.775143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T19:52:49.324500Z digest=sha256:7e6517f22e22b7e5bd27891f894f96e5de8a9dd110dd54daf61ba7b8b5380d20

Observation d81c616b-3c83-48c7-8893-5aa8fa3fb860 · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:08:26.182245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:da17252e2443b49d22f187f494b32c466441c4fa9ae8ce5505e59e5997f6c122

Observation ae7d3cc9-89ac-4a6d-bc98-253c89ec33fc · inbound

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions cites this paper.

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.331881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T05:49:57.957009Z digest=sha256:6bb26d18147771e31eab62c35c30734706a5a5605c6af26574d7942ca3dd7857

Observation 4524c26e-f802-4aa7-888b-6e408c19af54 · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:57.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:3faf9267f3c981e3699c931ba88ed57184e13d5903a3284a99538c72fd953c02

Observation 60f2ccbc-e987-460d-a2fc-12697d43e592 · inbound

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback cites this paper.

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:23:10.609159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T09:22:06.285118Z digest=sha256:d8b47e1b5d371ac98f6a069a761d58a42271ed49aa053004466300b1ce905b69

Observation 7d49d5d0-ece7-4be7-bb7e-e28556ea6543 · inbound

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models cites this paper.

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.672803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T06:25:56.404246Z digest=sha256:80dc846823be10e61e394d1f69f8de6a103cdbb2e2c159901c1196df871fdd0c

Observation 9e3c6695-b408-4083-a1cc-6c9803c17d38 · inbound

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy cites this paper.

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T04:43:45.592808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:43:45.592808Z digest=sha256:054fa16951ec86ca7b133395e8f4b1a2bdbea00aede98edf58cad8bdcd9a557d