Pith. sign in

Paper Citation Record · LEDGER

Are Large Language Models Memorizing Bug Benchmarks?

As of 15 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2411.13323.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13323 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:42.962606Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:20:21.052795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T00:16:15.729268Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19cccd37-eae3-4fe3-9883-69a6b40c60dc · outbound

This paper cites A survey on software fault localization,.

Are Large Language Models Memorizing Bug Benchmarks? A survey on software fault localization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.177178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.252494Z digest=sha256:c281ade62caf5470f0672b62d0d0f4621df1ddb06bbac8f6fbeb48378dc78ce7

Observation 619004af-2b05-49c6-aad8-7356b5ef0525 · outbound

This paper cites Automated Program Repair,.

Are Large Language Models Memorizing Bug Benchmarks? Automated Program Repair,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.163595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.256974Z digest=sha256:bfb66914ad5d555fbfc083ccd350f69a1891f765b38dde870edd01212080cfc2

Observation 9546f144-40e9-40a9-be7a-788c08e8949b · outbound

This paper cites Defects4j: a database of existing faults to enable controlled testing studies for java programs,.

Are Large Language Models Memorizing Bug Benchmarks? Defects4j: a database of existing faults to enable controlled testing studies for java programs,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.151216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.261686Z digest=sha256:ddffa2f552dbe5c810da74707c967b4279ed65f48d3285ee4f1bb33a2cc01433

Observation a52e983f-e0fe-4a17-b62e-0a3a1e26410c · outbound

This paper cites Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,.

Are Large Language Models Memorizing Bug Benchmarks? Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.139126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.265987Z digest=sha256:f8e5ef246c0e3df124386a5311ad1d11a9084b5d1a65d8ce27503bd8932e26d4

Observation d0bfa93d-fe52-4f03-8c37-1e0cf92ad63f · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues?.

Are Large Language Models Memorizing Bug Benchmarks? SWE-bench: Can language models resolve real-world github issues?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.124917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.270348Z digest=sha256:cc48354f1832e13893138f675a5bc02502c828e9bd9a9431ced18b08d08614bb

Observation 05131f45-ffc5-4890-a61f-3b941cf05c6b · outbound

This paper cites Concerned with Data Contamination? Assessing Countermeasures in Code Language Model.

Are Large Language Models Memorizing Bug Benchmarks? Concerned with Data Contamination? Assessing Countermeasures in Code Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.274919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.274919Z digest=sha256:fdd02684512e5c37e62880d93eea239b2826673f0d26ae8618e12c7e8d49ae9d

Observation e1c4d370-6e1e-4cf5-bcab-de488e646eed · outbound

This paper cites Leakage and the Reproducibility Crisis in ML-based Science.

Are Large Language Models Memorizing Bug Benchmarks? Leakage and the Reproducibility Crisis in ML-based Science

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.280170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.280170Z digest=sha256:c979f5e2f9844f051d6fd4cbc6ec8c3394bc0ceb8b07de56e50b571d4099f3ae

Observation 91903aa1-6de5-4571-be83-ce728ec2225b · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Are Large Language Models Memorizing Bug Benchmarks? Benchmarking Benchmark Leakage in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.284289Z digest=sha256:38ade4985e7b72b162dd7c554416ccd06f2319e0e94b3d6a3bf0dfbfc0b8d5e1

Observation 4f9e806b-50c5-4845-a892-601a9d8ad91c · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.288484Z digest=sha256:ce44a04cfec390819d782eb511956486d89b7a44bccfc8f00a91ecfaae665879

Observation 9cdb5a41-7298-4f92-9321-46d76c4fce2f · outbound

This paper cites Bugsc++: A highly usable real world defect benchmark for c/c++,.

Are Large Language Models Memorizing Bug Benchmarks? Bugsc++: A highly usable real world defect benchmark for c/c++,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.296648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.296648Z digest=sha256:590e7d8c8e5a3b0e670da5df8eba2474b3df6da7b36314e7eb71f7c8c19d7526

Observation d09c4910-7504-420a-9664-9fc61b9e9c9c · outbound

This paper cites Gitbug-java: A reproducible benchmark of recent java bugs,.

Are Large Language Models Memorizing Bug Benchmarks? Gitbug-java: A reproducible benchmark of recent java bugs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.959940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.300069Z digest=sha256:98e0e06f591dc92f283591d06120bff88120c5525cf825ad8f5562f4ca872aa1

Observation 7fb3822f-891b-4631-97b0-6ce0bed8fcee · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Are Large Language Models Memorizing Bug Benchmarks? Agentless: Demystifying LLM-based Software Engineering Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.304045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.304045Z digest=sha256:ac7ca84eae2c4275f97fba352267eb8747138e12e8e0492cab15cdae39c1d540

Observation 9db61581-7331-4858-9b14-7e0d47a8eb4c · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Are Large Language Models Memorizing Bug Benchmarks? OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.308840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.308840Z digest=sha256:39923dd8681d59f3a8ee4236eaf783f6c8940366dc25fe87c73b7d63186e1427

Observation 734c1bc4-4623-40cb-aa05-49b7d4d376e7 · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

Are Large Language Models Memorizing Bug Benchmarks? SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.349704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.349704Z digest=sha256:850de3f566daffe74aae09ffe4b82e5da4cf0e2f2aba690a2bf0f120fb1eece6

Observation c6e344e4-9bad-4ae1-9a67-b4b9258ae866 · outbound

This paper cites On the resemblance and containment of documents,.

Are Large Language Models Memorizing Bug Benchmarks? On the resemblance and containment of documents,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.877092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.405481Z digest=sha256:3c23d7d98ab88a7fcf38566a34eb43d916f9e5baca7420a1483367154ac9a363

Observation 6242c7ba-190b-4661-ac2c-3fde315d8ec7 · outbound

This paper cites Similarity search in high dimensions via hashing,.

Are Large Language Models Memorizing Bug Benchmarks? Similarity search in high dimensions via hashing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.863930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.506527Z digest=sha256:818bf646fd469043dcb550167fec39131ea1b5ab74c66c9f794e5c3cf0a4519d

Observation 5233ba2b-6590-4991-bb27-c1ab79d191d5 · outbound

This paper cites Codegen: An open large language model for code with multi-turn program synthesis,.

Are Large Language Models Memorizing Bug Benchmarks? Codegen: An open large language model for code with multi-turn program synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.852055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.510429Z digest=sha256:eb41bd097242f9ed47fea6b3ca379cf4f6eb84a2d637f3b5139ea69b480fac1c

Observation fcf00b6b-9882-4094-a865-574eb4dff3d6 · outbound

This paper cites Code llama: Open foundation models for code,.

Are Large Language Models Memorizing Bug Benchmarks? Code llama: Open foundation models for code,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.837616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.514409Z digest=sha256:12a551280d2aadc311baed40debdae4f817eecfcebdc4a86c2bb0c6c425d982b

Observation 5722db26-a0a7-4e0f-b402-92c10b750856 · outbound

This paper cites Llama: Open and efficient foundation language models,.

Are Large Language Models Memorizing Bug Benchmarks? Llama: Open and efficient foundation language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.522310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.522310Z digest=sha256:15c02c4aee4056294d513c183c097af6053d67a8801cbe7e03344162dbbe78a2

Observation 6b7abd74-868e-4562-a69b-ed465abbc3a7 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? Code Llama: Open Foundation Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.518109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.518109Z digest=sha256:94bf601a9172b614d16c2f4b3078fcf20a1b87c82c1e469c650e9c359526a176

Observation 5648e50f-7837-4773-8691-f79a815bf547 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Are Large Language Models Memorizing Bug Benchmarks? Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.530024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.530024Z digest=sha256:a08bcd8934c8493e183fff65c700c53448fb32702df40496bf9c48167ade43c7

Observation 6a8e0c35-14a6-47d8-9314-ecbfd02df6bf · outbound

This paper cites Starcoder 2 and the stack v2: The next generation,.

Are Large Language Models Memorizing Bug Benchmarks? Starcoder 2 and the stack v2: The next generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.776154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.525883Z digest=sha256:8b106b6f52651322ad8d6664b3fa3afe0e138178a46646bff8805200544aef49

Observation c97a4665-f0e4-47f3-a172-c052093eb1fa · outbound

This paper cites Mistral 7B.

Are Large Language Models Memorizing Bug Benchmarks? Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.758008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.758008Z digest=sha256:08342fb07956b2643493022543a90e8800b9fa71239cc701977f9c1a0d731c92

Observation 29f57848-3c77-446d-ac1a-2e9bf18ee480 · outbound

This paper cites Codegemma: Open code models based on gemma,.

Are Large Language Models Memorizing Bug Benchmarks? Codegemma: Open code models based on gemma,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.686661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.586465Z digest=sha256:87580e74c12155790dfcae3bc8afb0b848ea53bdf2c297a57cfa3c6233a2c7f6

Observation 5f6b8255-4788-4973-bdc5-e68bb0434e98 · outbound

This paper cites Repairllama: Efficient representations and fine-tuned adapters for program repair,.

Are Large Language Models Memorizing Bug Benchmarks? Repairllama: Efficient representations and fine-tuned adapters for program repair,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.660234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.795695Z digest=sha256:956a89e7a6396398c854c7049d5e06d7525347fcd2004290fdf087a8ccd16557

Observation 1a7cc93e-7ad8-43a2-8bcb-3823f9b331c6 · outbound

This paper cites Large language model for vulnerability detection: Emerging results and future directions,.

Are Large Language Models Memorizing Bug Benchmarks? Large language model for vulnerability detection: Emerging results and future directions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.647047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.804097Z digest=sha256:c45cc82fc1f7addbc04e8f1794e90868d35bd56e8010f51b8f24230710ce259e

Observation 10514c65-8456-4621-9750-548c46acfb89 · outbound

This paper cites Large language models for test-free fault localization,.

Are Large Language Models Memorizing Bug Benchmarks? Large language models for test-free fault localization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.673850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.791692Z digest=sha256:e57de52d7c97af7d926da0d554999a4a48acb4c8c079b2018474740844cf2b52

Observation 51d34f47-928c-4ef0-b6c9-2fea25c1c71e · outbound

This paper cites A study of the uniqueness of source code,.

Are Large Language Models Memorizing Bug Benchmarks? A study of the uniqueness of source code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.468959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.811382Z digest=sha256:a25dbcc124a02e90389bc54f4fde55d4e577339cad888057d0eee2c1be5fb686

Observation 59e1695b-2a69-4bf8-ab28-b7dafbd097c9 · outbound

This paper cites RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair.

Are Large Language Models Memorizing Bug Benchmarks? RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.799148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.799148Z digest=sha256:cedb08fef5375c0567e002abd8e2419de3f55c5dee601ea98ac111743380db5f

Observation 49ac9234-caa2-4cd3-aa83-76a8f17f36c4 · outbound

This paper cites Memorization without overfitting: Analyzing the training dynamics of large language models,.

Are Large Language Models Memorizing Bug Benchmarks? Memorization without overfitting: Analyzing the training dynamics of large language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.455722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.819025Z digest=sha256:b4d7cff56751a61e8e77482f38187d89e34ae988116c5bf156f0e7139438e046

Observation 41c367a4-fe6c-4c5f-af46-ae0e0176eb1d · outbound

This paper cites The stack: 3 TB of permissively licensed source code,.

Are Large Language Models Memorizing Bug Benchmarks? The stack: 3 TB of permissively licensed source code,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.575402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.807962Z digest=sha256:7371e52cf5ff2552f5398fca4d0f8a8045fab0bf0bdb8e84f483e38681a7f92f

Observation 5b037e62-a0b0-4319-9405-f770df87ecad · outbound

This paper cites Measuring massive multitask language understanding,.

Are Large Language Models Memorizing Bug Benchmarks? Measuring massive multitask language understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.440137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.826381Z digest=sha256:c674d8e7a991d215be6400563db4def4fc69e00283e642f8f817066584429eb3

Observation aaaadc1c-7b6a-4bde-b9b6-f7a4fdaaff5e · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Are Large Language Models Memorizing Bug Benchmarks? An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.814911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.814911Z digest=sha256:3ab0568c3f51f233cb11fef0db9910ebb5622298fbaf6acfe3bff71c5ab13cab

Observation ccc78991-03b9-4c7d-9762-de97282f6fd7 · outbound

This paper cites Program Synthesis with Large Language Models.

Are Large Language Models Memorizing Bug Benchmarks? Program Synthesis with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.838268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.838268Z digest=sha256:0b86ccba0922925ccbee0bb578061b104c9a076aa3329fbbee4656b61b5b79dc

Observation a1acba78-afc4-4fa8-8b08-09bdf13a6c33 · outbound

This paper cites Squad: 100, 000+ questions for machine comprehension of text,.

Are Large Language Models Memorizing Bug Benchmarks? Squad: 100, 000+ questions for machine comprehension of text,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.822872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.822872Z digest=sha256:e1e89eedac2cfecdfc0b5949ecee02d88e4981288b354ff5efac5b664c63bbc4

Observation 7d5cf801-2987-42b0-9aff-464a161cbd97 · outbound

This paper cites Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT.

Are Large Language Models Memorizing Bug Benchmarks? Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.845327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.845327Z digest=sha256:c4da60cd345015d8234bdcfc4f07b8075347bdefa4e412c46c031d2cc1490cfc

Observation a3a57390-4812-4472-819f-e5750cd47e0a · outbound

This paper cites Evaluating large language models trained on code,.

Are Large Language Models Memorizing Bug Benchmarks? Evaluating large language models trained on code,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.830570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.830570Z digest=sha256:1e559fe4852aface20a83105c2e29265121701a87b4f341a639351103e2abc3b

Observation b293650d-04a7-415b-9989-f12de4f16fe7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.893040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.893040Z digest=sha256:b4b708711e948e9dca9646c425798df57a257dfd563fd1bb44a620d4207d0831

Observation 61797f08-cfb1-4f59-b0fd-409ea0d30555 · outbound

This paper cites Why does your data leak? uncovering the data leakage in cloud from mobile apps,.

Are Large Language Models Memorizing Bug Benchmarks? Why does your data leak? uncovering the data leakage in cloud from mobile apps,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.353535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.962606Z digest=sha256:282c4ce4cf0e44f79a833b4e855c2675c4ac8e52a2ced08f4fa592f21f1561a8

Observation 47de63b2-7354-4d2d-944e-93e212e32bf4 · outbound

This paper cites Measuring coding challenge competence with APPS,.

Are Large Language Models Memorizing Bug Benchmarks? Measuring coding challenge competence with APPS,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.418581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.841580Z digest=sha256:864be5464d3a4a71499166101d296ae1022c789d4e30a2fad862d9ae8e18c6e4

Observation aa29b1cc-1416-4b5b-b1ef-8677991695f3 · outbound

This paper cites Detecting pretraining data from large language models,.

Are Large Language Models Memorizing Bug Benchmarks? Detecting pretraining data from large language models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.405143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:42.849344Z digest=sha256:765146877067d836614765bc62f6ddd78963141b2e28e66ba3f4dfd9d6dcf187

Observation 3b338049-eaca-4c1a-bed3-2f8252848170 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.958319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.958319Z digest=sha256:07a90573e5d5fbe6bf8c846a04e8350c1a76fe3ea88ab9d4c8ae36a00915e9ec

Observation 371f832b-16fd-45fa-8db1-91e2625b18a2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Are Large Language Models Memorizing Bug Benchmarks? Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.833805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.833805Z digest=sha256:e3c86d26a6cb6985b4ca2bf8a423ceac80b9d7e4e8f6577ba8eebd0363972895

Observation 860f927a-9d1b-4c35-b4fc-9d5cea66c2cc · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.292215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.292215Z digest=sha256:cb4bdd94d36feb36ee50efabf4a06516ac3210e163a24f338f9aea1d4ab26b5e

Observation df50b642-a106-4b59-b337-7ca5e5baebda · outbound

This paper cites CodeGemma: Open Code Models Based on Gemma.

Are Large Language Models Memorizing Bug Benchmarks? CodeGemma: Open Code Models Based on Gemma

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.656619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.656619Z digest=sha256:0279965c6118ad0222700bfc0050a01403e8dec03a8ada552219b71d4f910940

Pith citing papers

Observation 904278f3-d300-46fa-94f9-2b1ea3fc9164 · inbound

CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis cites this paper.

CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis Are Large Language Models Memorizing Bug Benchmarks?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:20:21.116359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:20:21.052795Z digest=sha256:b29bc45892a72c194f8202814a4bcf65427924c862ad6eeea9b0df60f98b2f98