Pith. sign in

Paper Citation Record · LEDGER

Addressing Data Leakage in HumanEval Using Combinatorial Test Design

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.01526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01526 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:20:21.524087Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36a47dbc-3bde-46d7-bf17-f69ce5f5ccef · outbound

This paper cites Large language models for software engineering: Sur- vey and open problems,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Large language models for software engineering: Sur- vey and open problems,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.759238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.426420Z digest=sha256:7b74bc4cf9b51ea247b4f9af53fc68ca226943c4dfa826cdab70a202ab0dd780

Observation 0dba7e71-d769-4c94-987a-01f9d7561e66 · outbound

This paper cites Software testing with large language models: Survey, landscape, and vision,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Software testing with large language models: Survey, landscape, and vision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.431387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.431387Z digest=sha256:24941ae656b7fe8aca5233272ef91b49c43e872e36626985e7499b4526f27dfd

Observation 0012f093-1f51-4be2-94c7-f904e3847ed0 · outbound

This paper cites Large language models for software engineering: A systematic literature review,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Large language models for software engineering: A systematic literature review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.435724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.435724Z digest=sha256:92a933c7b831098557f05ec313fa052aa92a7f26e5ca31f5e5f916e125afce4b

Observation f48b6488-b426-477f-bc63-905a50902254 · outbound

This paper cites A Survey on Evaluation of Large Language Models,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design A Survey on Evaluation of Large Language Models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.440549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.440549Z digest=sha256:a91d6a7a1bcaaf4bcf5ff8f9fa95786d0d509456020af70aa886e92149e96a07

Observation 068aba4a-d1a9-4192-8403-7350dcc76267 · outbound

This paper cites A Software Engineering Perspective on Testing Large Language Models: Research, Practice, Tools and Benchmarks.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design A Software Engineering Perspective on Testing Large Language Models: Research, Practice, Tools and Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.445014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.445014Z digest=sha256:0acc8019c7da450cb96dd2deed076e245f261a2e9f6e498780b67c7ecca2971b

Observation 4555590f-4a22-452e-84b2-8e682f5bb20f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.449441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.449441Z digest=sha256:2a53011d19895f465ad6ea18677894fcabec5248a468a75538584277debc1ecc

Observation 236712a5-0d82-495a-a02b-1859d596513c · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Codereval: A benchmark of pragmatic code generation with generative pre-trained models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.454253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.454253Z digest=sha256:3a1203308ad11570d613006334c6669a3b81299bec0857babaab47c5bcbf9328

Observation 676a281d-38e1-4338-a028-8d24e3328ef6 · outbound

This paper cites Towards standarized benchmarks of LLMs in software modeling tasks: a conceptual framework,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Towards standarized benchmarks of LLMs in software modeling tasks: a conceptual framework,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.464228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.464228Z digest=sha256:05757eb7495f58f7f7449b3d0c012c64defe213cda51141c2b0a73dceb58932a

Observation 2b3d364f-f575-4ba1-9924-bdbbe3c5d0ad · outbound

This paper cites Using benchmarking infrastructure to evaluate llm performance on cs concept inventories: Challenges, opportunities, and critiques,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Using benchmarking infrastructure to evaluate llm performance on cs concept inventories: Challenges, opportunities, and critiques,

Reference 9

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.463999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.468618Z digest=sha256:286b30b7c26790eec021c5a6414e2c4848cbf65d2f357a0512a44d315af55858

Observation 149a2248-9216-4ffd-8f0b-ba279108ddea · outbound

This paper cites Condefects: A complementary dataset to address the data leakage concern for llm-based fault localization and program repair,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Condefects: A complementary dataset to address the data leakage concern for llm-based fault localization and program repair,

Reference 10

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.285490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.472665Z digest=sha256:c6a7876a9197149aa4f4725e631c70c6b74c4b518e0c9b8e19cd3ec761489850

Observation d67b6f95-605e-4ca9-836a-4789e64b2353 · outbound

This paper cites BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation,

Reference 11

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.096588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.476747Z digest=sha256:6dbe3ab1f519d83d45480f24e166712d02f71a9410db1e86ece58d2909353ac5

Observation ffd2df8d-b59f-4695-9cc1-eb5b565deece · outbound

This paper cites Holistic Evaluation of Language Models.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.483043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.483043Z digest=sha256:f8187751b0aa9a819ef4e109ad1bfc35d5235826844596174b171e5a34a77035

Observation 5e8413bb-b11a-41b9-b492-e82ad20b5500 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.487565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.487565Z digest=sha256:8af8e1725e115c70f1b0622f340e877a7d4404cffce679628769d816f39c881e

Observation 06b07fee-2001-47b6-9caf-33dbff040a34 · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.730332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.491898Z digest=sha256:c9968879ffcd7a4ec993223e70470d61975e67e79e5201a5210d16daec343b65

Observation d241f699-5b41-4b59-9583-b9b5290bfc30 · outbound

This paper cites Task Contamination: Language Models May Not Be Few-Shot Anymore,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Task Contamination: Language Models May Not Be Few-Shot Anymore,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.717836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.495785Z digest=sha256:9843ea77823d3f2dfcf5abf7f49a6756f3deed9532b6c46130df40207e911558

Observation 06bc2373-2751-4211-9c62-f2b9a7b2ea26 · outbound

This paper cites A test generation strategy for pairwise testing,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design A test generation strategy for pairwise testing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.705319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.499796Z digest=sha256:eaf2c97ee40c0d5df13b14a90d929e6797ed605064b595d8625ef43191559962

Observation 42dce237-fb1f-419d-beb3-d164c71fa7ce · outbound

This paper cites Combinatorial test design in practice,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Combinatorial test design in practice,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.692897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.504311Z digest=sha256:e50e022fdbc6d1334a9043bdd4248021ab976c170dd8d80ca25406dc4177f595

Observation 9944ad66-216d-4120-8a31-196d1054daef · outbound

This paper cites Using benchmarking to advance research: a challenge to software engineering,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Using benchmarking to advance research: a challenge to software engineering,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-12T04:20:21.926313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.508384Z digest=sha256:46b862fba4e6ec1724b3d043a1feb24bc5d4663355d23c4dbfbc6bd561c76e42

Observation babcd247-2587-4e4a-bdec-24a7767e85fd · outbound

This paper cites Using combinatorial benchmark construction to improve the assessment of concurrency bug detection tools,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Using combinatorial benchmark construction to improve the assessment of concurrency bug detection tools,

Reference 19

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:21.789763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.512271Z digest=sha256:848c1126fdd7e0272ac320e35db0292066147e79d8fdbe4466589dc77f161617

Observation 49c0fa33-9bfe-4b02-a60e-1b8da0d396f9 · outbound

This paper cites Applying pairwise combinatorial testing to large language model testing,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Applying pairwise combinatorial testing to large language model testing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.679251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.516299Z digest=sha256:688f159fb5158cf5f858ab78e35b94b0d0aa83ddb2bf255d2806572a76c2d506

Observation a62219f4-63e0-4c44-b634-98e80ad385fb · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.520097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.520097Z digest=sha256:b0ca93b79d7373196d2c7497935df98d9771f0bb00f7158a418aef3be3e61245

Observation 41996491-cd86-4db3-8569-cc51fdad9ac3 · outbound

This paper cites Dynamic Evaluation of Large Language Models by Meta Probing Agents.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.524087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.524087Z digest=sha256:86fca7f6cfbcc635543de840b4a720b18f0de6f2899dc2b5f0157d1571ed74cb

Observation 8e08a151-0be0-48f5-9029-1100fa38c204 · outbound

This paper cites Available: https://doi.org/10.1145/3597503.3623316.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Available: https://doi.org/10.1145/3597503.3623316

Reference 2024

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.642217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:20:21.459075Z digest=sha256:647b2a12b3664371226902f4b55185d5b2f4689b7f8e3a2098e9018218e3cfe4

Pith citing papers

No inbound Pith citation observations are available.