Pith. sign in

Paper Citation Record · LEDGER

Are Large Language Models Memorizing Bug Benchmarks?

As of 15 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2411.13323.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13323 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:42.962606Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:20:21.052795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T00:16:15.729268Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19cccd37-eae3-4fe3-9883-69a6b40c60dc · outbound

This paper cites A survey on software fault localization,.

Are Large Language Models Memorizing Bug Benchmarks? A survey on software fault localization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.177178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.252494Z digest=sha256:20c98d2f8a442b955d77e8859828c776c224b4afb3fab78f50859c83b53f1c5f

Observation 619004af-2b05-49c6-aad8-7356b5ef0525 · outbound

This paper cites Automated Program Repair,.

Are Large Language Models Memorizing Bug Benchmarks? Automated Program Repair,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.163595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.256974Z digest=sha256:f78b10d0d36563e6e57b03f7c6a7e61d7d7824517acef53f49f5e7c8eb6bf3e5

Observation 9546f144-40e9-40a9-be7a-788c08e8949b · outbound

This paper cites Defects4j: a database of existing faults to enable controlled testing studies for java programs,.

Are Large Language Models Memorizing Bug Benchmarks? Defects4j: a database of existing faults to enable controlled testing studies for java programs,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.151216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.261686Z digest=sha256:fab2e01cbd7ad758c53a739d96b080567a9e9045ee0aef629a79956a6695a1f1

Observation a52e983f-e0fe-4a17-b62e-0a3a1e26410c · outbound

This paper cites Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,.

Are Large Language Models Memorizing Bug Benchmarks? Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.139126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.265987Z digest=sha256:7796bd9e1927d3e4580fdf55a4c7407b815ebf9063ca89d2378566c6bd953514

Observation d0bfa93d-fe52-4f03-8c37-1e0cf92ad63f · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues?.

Are Large Language Models Memorizing Bug Benchmarks? SWE-bench: Can language models resolve real-world github issues?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.124917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.270348Z digest=sha256:60c7542691fe22f723c11efdcc6e69d5e0a6c07b4be4615536e7c2ec377cd70d

Observation 05131f45-ffc5-4890-a61f-3b941cf05c6b · outbound

This paper cites Concerned with Data Contamination? Assessing Countermeasures in Code Language Model.

Are Large Language Models Memorizing Bug Benchmarks? Concerned with Data Contamination? Assessing Countermeasures in Code Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.274919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.274919Z digest=sha256:178295e13847be21484fc6d694d37e2de19513fb06ac6a78df86ae31a62e3557

Observation e1c4d370-6e1e-4cf5-bcab-de488e646eed · outbound

This paper cites Leakage and the Reproducibility Crisis in ML-based Science.

Are Large Language Models Memorizing Bug Benchmarks? Leakage and the Reproducibility Crisis in ML-based Science

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.280170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.280170Z digest=sha256:874305d8dc3e0c1576727592c3a337d7f320a96386fa80a124b1aa96ee61a0a5

Observation 91903aa1-6de5-4571-be83-ce728ec2225b · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Are Large Language Models Memorizing Bug Benchmarks? Benchmarking Benchmark Leakage in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.284289Z digest=sha256:c6fa70d34059b29d93e9e60e6ef609e5314901b9c4ef4dd9fb928ff308847326

Observation 4f9e806b-50c5-4845-a892-601a9d8ad91c · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.288484Z digest=sha256:f77bea45c3889ca081cfaffc49446d14a383cec3e25ee9d31738de92aa5ad253

Observation 9cdb5a41-7298-4f92-9321-46d76c4fce2f · outbound

This paper cites Bugsc++: A highly usable real world defect benchmark for c/c++,.

Are Large Language Models Memorizing Bug Benchmarks? Bugsc++: A highly usable real world defect benchmark for c/c++,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.296648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.296648Z digest=sha256:3c5310263226997ed37ebb8f7fee8ea0f52d744f0f15dfeef27ee455f02915cf

Observation d09c4910-7504-420a-9664-9fc61b9e9c9c · outbound

This paper cites Gitbug-java: A reproducible benchmark of recent java bugs,.

Are Large Language Models Memorizing Bug Benchmarks? Gitbug-java: A reproducible benchmark of recent java bugs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.959940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.300069Z digest=sha256:39bed2d311f9a24613af25db7305bae2e1f8b9237d46e0e7a525167edeb23300

Observation 7fb3822f-891b-4631-97b0-6ce0bed8fcee · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Are Large Language Models Memorizing Bug Benchmarks? Agentless: Demystifying LLM-based Software Engineering Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.304045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.304045Z digest=sha256:b4beaded90afb259aa0b5bbf5b2a9d5f09fab9f0bcf2e5173904b5de3f94ae05

Observation 9db61581-7331-4858-9b14-7e0d47a8eb4c · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Are Large Language Models Memorizing Bug Benchmarks? OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.308840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.308840Z digest=sha256:88aa5fe76a9f7980e9de4f53312e597f65429fcc8002689ea262b9cc5a5cd06c

Observation 734c1bc4-4623-40cb-aa05-49b7d4d376e7 · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

Are Large Language Models Memorizing Bug Benchmarks? SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.349704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.349704Z digest=sha256:7e843dc75fd23b7c99d40580270dc0b5528cdb656333eea4c2a20892126bc51b

Observation c6e344e4-9bad-4ae1-9a67-b4b9258ae866 · outbound

This paper cites On the resemblance and containment of documents,.

Are Large Language Models Memorizing Bug Benchmarks? On the resemblance and containment of documents,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.877092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.405481Z digest=sha256:c6408b4612865251744b8b373c9b5e66643dc397afaa7ee430042c9854619cf7

Observation 6242c7ba-190b-4661-ac2c-3fde315d8ec7 · outbound

This paper cites Similarity search in high dimensions via hashing,.

Are Large Language Models Memorizing Bug Benchmarks? Similarity search in high dimensions via hashing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.863930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.506527Z digest=sha256:35aeefa06aa4732ba017cf927190781fa1529015d109b3216845c6e8432a714e

Observation 5233ba2b-6590-4991-bb27-c1ab79d191d5 · outbound

This paper cites Codegen: An open large language model for code with multi-turn program synthesis,.

Are Large Language Models Memorizing Bug Benchmarks? Codegen: An open large language model for code with multi-turn program synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.852055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.510429Z digest=sha256:b31655adac0392a5afc5ba47b638ea0701b31f62d29cc64af433f419debe356a

Observation fcf00b6b-9882-4094-a865-574eb4dff3d6 · outbound

This paper cites Code llama: Open foundation models for code,.

Are Large Language Models Memorizing Bug Benchmarks? Code llama: Open foundation models for code,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.837616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.514409Z digest=sha256:d3a7eb563f2429be15935526a27d05b41635fca89659640321909f63d7612138

Observation 5722db26-a0a7-4e0f-b402-92c10b750856 · outbound

This paper cites Llama: Open and efficient foundation language models,.

Are Large Language Models Memorizing Bug Benchmarks? Llama: Open and efficient foundation language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.522310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.522310Z digest=sha256:5b5b73a1d8c54c73344b60be3cca57e82fb3812b846a0e274c5954a1725ca7d6

Observation 6b7abd74-868e-4562-a69b-ed465abbc3a7 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? Code Llama: Open Foundation Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.518109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.518109Z digest=sha256:8827b29c5c38b3046d19c1e9ae7c27984077e23a15684dcbb8d8b07120eeaa42

Observation 5648e50f-7837-4773-8691-f79a815bf547 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Are Large Language Models Memorizing Bug Benchmarks? Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.530024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.530024Z digest=sha256:e76bd5c23986793e65734ceed313fa1cc372785516ada9d62e12ff127fceeefd

Observation 6a8e0c35-14a6-47d8-9314-ecbfd02df6bf · outbound

This paper cites Starcoder 2 and the stack v2: The next generation,.

Are Large Language Models Memorizing Bug Benchmarks? Starcoder 2 and the stack v2: The next generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.776154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.525883Z digest=sha256:bdb7001f5ff7f9d56a0ce17c1951ae77e727d0013bcb4e149448b007fc791c71

Observation c97a4665-f0e4-47f3-a172-c052093eb1fa · outbound

This paper cites Mistral 7B.

Are Large Language Models Memorizing Bug Benchmarks? Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.758008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.758008Z digest=sha256:083b18cec3d133cb690a53d97b7581814aacc8cf83038d5c34138ec7908e9061

Observation 29f57848-3c77-446d-ac1a-2e9bf18ee480 · outbound

This paper cites Codegemma: Open code models based on gemma,.

Are Large Language Models Memorizing Bug Benchmarks? Codegemma: Open code models based on gemma,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.686661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.586465Z digest=sha256:71faaa2ff88eedd4e4c9ace6e5e45b000e729bfcad7d8f42bf34419c9e24f386

Observation 5f6b8255-4788-4973-bdc5-e68bb0434e98 · outbound

This paper cites Repairllama: Efficient representations and fine-tuned adapters for program repair,.

Are Large Language Models Memorizing Bug Benchmarks? Repairllama: Efficient representations and fine-tuned adapters for program repair,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.660234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.795695Z digest=sha256:c33519be3dd538fc60e1d44cfb3b8ffd472d3405f00f4695872e4b16369c9ed5

Observation 1a7cc93e-7ad8-43a2-8bcb-3823f9b331c6 · outbound

This paper cites Large language model for vulnerability detection: Emerging results and future directions,.

Are Large Language Models Memorizing Bug Benchmarks? Large language model for vulnerability detection: Emerging results and future directions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.647047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.804097Z digest=sha256:d9743de6cb09c633b6d856359a9ff30fba4122bd9dd970d60dbf9b8544fb7e61

Observation 10514c65-8456-4621-9750-548c46acfb89 · outbound

This paper cites Large language models for test-free fault localization,.

Are Large Language Models Memorizing Bug Benchmarks? Large language models for test-free fault localization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.673850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.791692Z digest=sha256:53fef449ed390364d6a4cdd5df5e61ca07417a7cdaab4f59d110515e37f8ed43

Observation 51d34f47-928c-4ef0-b6c9-2fea25c1c71e · outbound

This paper cites A study of the uniqueness of source code,.

Are Large Language Models Memorizing Bug Benchmarks? A study of the uniqueness of source code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.468959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.811382Z digest=sha256:ca0c438221ae01adf46c4a0128be3849bc53c864b19f48cf5d5cf83e753147e7

Observation 59e1695b-2a69-4bf8-ab28-b7dafbd097c9 · outbound

This paper cites RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair.

Are Large Language Models Memorizing Bug Benchmarks? RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.799148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.799148Z digest=sha256:f0a983d8f8ee63c6a75d70f017099681671433a1f772dd1ce667d06c5f9f31f1

Observation 49ac9234-caa2-4cd3-aa83-76a8f17f36c4 · outbound

This paper cites Memorization without overfitting: Analyzing the training dynamics of large language models,.

Are Large Language Models Memorizing Bug Benchmarks? Memorization without overfitting: Analyzing the training dynamics of large language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.455722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.819025Z digest=sha256:904855ee9d053197802868f72456dd132768736d392e26c18a8fad9110ba505c

Observation 41c367a4-fe6c-4c5f-af46-ae0e0176eb1d · outbound

This paper cites The stack: 3 TB of permissively licensed source code,.

Are Large Language Models Memorizing Bug Benchmarks? The stack: 3 TB of permissively licensed source code,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.575402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.807962Z digest=sha256:ee6240949b78f7cd5c8fb32ccdad0830a5580e08ec8f59cae89843ebe44801de

Observation 5b037e62-a0b0-4319-9405-f770df87ecad · outbound

This paper cites Measuring massive multitask language understanding,.

Are Large Language Models Memorizing Bug Benchmarks? Measuring massive multitask language understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.440137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.826381Z digest=sha256:3565db13e3d818e01f4c2df76048c29abc3056a3a9212f32f060f7fb27572561

Observation aaaadc1c-7b6a-4bde-b9b6-f7a4fdaaff5e · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Are Large Language Models Memorizing Bug Benchmarks? An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.814911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.814911Z digest=sha256:ac60395e46d77fdf31d02adf17cb6601ce76bc021addc3ee0a5d5b4e26c23ba6

Observation ccc78991-03b9-4c7d-9762-de97282f6fd7 · outbound

This paper cites Program Synthesis with Large Language Models.

Are Large Language Models Memorizing Bug Benchmarks? Program Synthesis with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.838268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.838268Z digest=sha256:4107d810e98bcffe11b073487b9597158e7d7c016198ff65a163d1deb35c2bb1

Observation a1acba78-afc4-4fa8-8b08-09bdf13a6c33 · outbound

This paper cites Squad: 100, 000+ questions for machine comprehension of text,.

Are Large Language Models Memorizing Bug Benchmarks? Squad: 100, 000+ questions for machine comprehension of text,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.822872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.822872Z digest=sha256:172ddf1883fcd60553b12930b5ebf3df6d1db1f7b0650cfc9637e68256b96e60

Observation 7d5cf801-2987-42b0-9aff-464a161cbd97 · outbound

This paper cites Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT.

Are Large Language Models Memorizing Bug Benchmarks? Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.845327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.845327Z digest=sha256:e4fe0ebed578a510f8c90087161ffbc083f7b8e0d75a5e90ee3235abb42cdd98

Observation a3a57390-4812-4472-819f-e5750cd47e0a · outbound

This paper cites Evaluating large language models trained on code,.

Are Large Language Models Memorizing Bug Benchmarks? Evaluating large language models trained on code,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.830570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.830570Z digest=sha256:a6ea222e2cb7b62c57615ce39bf68ecadc54ce54abe71254f854ac086313a6d5

Observation b293650d-04a7-415b-9989-f12de4f16fe7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.893040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.893040Z digest=sha256:91c4628acfc713038671f69042d197f28f5ac94c569f86ca02f07809fc36afe5

Observation 61797f08-cfb1-4f59-b0fd-409ea0d30555 · outbound

This paper cites Why does your data leak? uncovering the data leakage in cloud from mobile apps,.

Are Large Language Models Memorizing Bug Benchmarks? Why does your data leak? uncovering the data leakage in cloud from mobile apps,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.353535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.962606Z digest=sha256:87e1815559a279aafa909bc3fb280b487daa11129f3688f6f842d8621c882a52

Observation 47de63b2-7354-4d2d-944e-93e212e32bf4 · outbound

This paper cites Measuring coding challenge competence with APPS,.

Are Large Language Models Memorizing Bug Benchmarks? Measuring coding challenge competence with APPS,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.418581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.841580Z digest=sha256:db5e6743b74650b48ed624cf99404e57b16c4874085dd17dfc3bc6e83ccd0d17

Observation aa29b1cc-1416-4b5b-b1ef-8677991695f3 · outbound

This paper cites Detecting pretraining data from large language models,.

Are Large Language Models Memorizing Bug Benchmarks? Detecting pretraining data from large language models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.405143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:38:42.849344Z digest=sha256:08ed47a3190b4322c8c5114709ccd3fa4fb834d8e5e0ce793169d7c59200f979

Observation 3b338049-eaca-4c1a-bed3-2f8252848170 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.958319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.958319Z digest=sha256:9a93a9026955ab5141d62a953c3882b0980e9f67b1c4759646fff7fcbe40b290

Observation 371f832b-16fd-45fa-8db1-91e2625b18a2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Are Large Language Models Memorizing Bug Benchmarks? Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.833805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.833805Z digest=sha256:da0944c32387e79fbd5f2d6e8f119331a5d8fbeae121a3b52a686cdb6869b59f

Observation 860f927a-9d1b-4c35-b4fc-9d5cea66c2cc · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.292215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.292215Z digest=sha256:df461e3f31a1da8b6003fc8e85404ee2794ce6640045f6165bf73080f66e8ebc

Observation df50b642-a106-4b59-b337-7ca5e5baebda · outbound

This paper cites CodeGemma: Open Code Models Based on Gemma.

Are Large Language Models Memorizing Bug Benchmarks? CodeGemma: Open Code Models Based on Gemma

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.656619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.656619Z digest=sha256:1f147b9960c931e4378a5ba2fdccb1c1afbbf810ebbaf49e9fbbd036ba0c70e8

Pith citing papers

Observation 904278f3-d300-46fa-94f9-2b1ea3fc9164 · inbound

CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis cites this paper.

CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis Are Large Language Models Memorizing Bug Benchmarks?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:20:21.116359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:20:21.052795Z digest=sha256:6675241cd42d6c4b21f0f8d108201457a3dc20911c1c32ab5ac03b6495734b6e