Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:37.864289Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 11 inbound Pith citation observations for arXiv:2412.21199.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:37.864289Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:46:23.589246Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T14:33:31.671093Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d876dc4f-c33c-4ac7-accd-52835c08fa0a · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d88c956-7f5e-47f4-9ad4-f0a8c7f1f6ec · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 080a5760-a8d3-474a-b9cc-8212ebbedc5f · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f8053017-7e59-4e98-8250-3d4326e88a78 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033749d6-46d4-4338-83ba-565e6f3dc325 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Multi-lingual Evaluation of Code Generation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 081d3982-20cd-4ed4-b84d-6b6827ac32a0 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Program Synthesis with Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c353f6-ca3b-4472-8e11-8c009960e6c4 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation A parallel corpus of Python functions and documentation strings for automated code documentation and code generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a4491f1-207e-4155-8203-e5941a53d6c1 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5be030-5340-4c47-aee8-273f73006960 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e24e4516-ad1c-4218-9688-cfc5990073b4 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7e684a6e-11d6-4abd-96f8-514f5772a171 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26248b7-fc10-45f4-97c7-8c4562ad2f08 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cf1cb287-f0a1-4222-bcd3-33b5898fb3d1 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation CoDesc: A Large Code-Description Parallel Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce31397-d5bc-4d4f-a982-7e2798b26cba · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977957ed-a753-4bde-8f22-ffae09f7d481 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Qwen2.5-Coder Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d5afda-6d22-456c-b57f-bf05192c3346 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4546af62-288e-4803-a79f-764d74d0ae2a · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Impact of Code Language Models on Automated Program Repair
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e313ef28-8728-4a80-9525-5fa9c17d9f06 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019ef2cc-36e5-42db-8d4b-fc9d237e9f18 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation InferFix: End-to-End Program Repair with LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d898fb23-5cf4-43b7-b2ef-a5be47bca329 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d659fbe8-440b-44d8-808f-9e09ed6ac76b · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation StarCoder: may the source be with you!
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8426c663-bc41-40d2-a51c-8d9638c55351 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6b7744f3-bc8a-4b86-8175-da019dd24d4b · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020fc0de-60e9-4288-9030-e7c11452b449 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9546eaaf-9d9c-4d2b-a67e-48fa0fdf7093 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 854ad207-6b18-4fbf-a254-ddb76d19f573 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4169652-c38d-4b82-988e-aa3ec49d70b1 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1748b129-b0bb-4a33-8323-a46721ac04b1 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51a1e376-df4a-47c1-b7e8-86bd9ef400f6 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6fea44e9-578c-491a-8da9-d674ba2718c3 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97a2495-8882-4308-99fc-a8ab6bf6a230 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f25c5153-1218-4f6d-a3f0-5ed636b533a0 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Code Llama: Open Foundation Models for Code
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d931f9-994f-445d-85ad-cbe24d7a93f6 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d639f24a-63f0-48b2-8fca-fa2fc2f95657 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4492f096-a198-4925-9dad-596d10ecb4d6 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0250df47-431e-4736-9b2c-6263f50e7bdb · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 66e6bfa1-eb77-495b-b6a6-c583ae4c3343 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Practical Program Repair in the Era of Large Pre-trained Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46bf6f97-003b-4d22-95f6-16f2c6ce59d3 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cb009c-685a-42c3-8b61-3237d64ddafe · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8abdda8-005d-44d1-a10c-884343b3e9e5 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15370ef-e38b-4f6e-962c-f751d784b109 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8c92bb-49a5-4c0b-af41-af0d143da344 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb3ab70-438b-4c3a-81d7-be9d5ef433e4 · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdb5c8e-c31b-4ff2-9365-c2e259feb06c · outbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0188b4d-f77f-400f-b809-c542f336ef2e · inbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1dc0812-c573-420c-a324-23f2aca5ef02 · inbound
Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e3b6f0-bf66-4c79-a1bc-9e4cc155d542 · inbound
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · inbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9a34e2-15bc-46cd-893c-47c92f4f07ed · inbound
LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d81c616b-3c83-48c7-8893-5aa8fa3fb860 · inbound
Agentic Frameworks for Reasoning Tasks: An Empirical Study HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ae7d3cc9-89ac-4a6d-bc98-253c89ec33fc · inbound
Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4524c26e-f802-4aa7-888b-6e408c19af54 · inbound
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 60f2ccbc-e987-460d-a2fc-12697d43e592 · inbound
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7d49d5d0-ece7-4be7-bb7e-e28556ea6543 · inbound
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e3c6695-b408-4083-a1cc-6c9803c17d38 · inbound
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.