Pith. sign in

Paper Citation Record · LEDGER

Extracting Training Data from Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2012.07805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.07805 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:37:25.072025Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

275
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f249adc-febb-4812-a104-ac684d731332 · inbound

The Pile: An 800GB Dataset of Diverse Text for Language Modeling cites this paper.

The Pile: An 800GB Dataset of Diverse Text for Language Modeling Extracting Training Data from Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:35:18.749655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T21:35:18.513342Z digest=sha256:936f0861e500d066e97cf9162077ce1beaba939d608a52b01d2fb9a9923e4011

Observation a001a767-4dd2-40a8-909e-e78e492c4507 · inbound

Deduplicating Training Data Makes Language Models Better cites this paper.

Deduplicating Training Data Makes Language Models Better Extracting Training Data from Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T13:39:31.714001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T13:36:55.210708Z digest=sha256:938b0396db7dc887143220dbafe4f5e7c95f8342d51d8557791bcbe7e887aa5b

Observation 5d360aae-b923-4573-b2ad-60a12b397e5b · inbound

Ethical and social risks of harm from Language Models cites this paper.

Ethical and social risks of harm from Language Models Extracting Training Data from Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:24:29.952856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T18:24:28.835688Z digest=sha256:b0529dd6fd27b487a73c6d6851fae06f9b138212af53f2d68160bbcfe47249be

Observation 3da7c7d4-035e-4215-a99a-aa3bd6df49c2 · inbound

LaMDA: Language Models for Dialog Applications cites this paper.

LaMDA: Language Models for Dialog Applications Extracting Training Data from Large Language Models

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:17:32.518336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:17:32.272353Z digest=sha256:3d66e574682846707cf518f363130e22b3f32e3810450dea337c633a167c353f

Observation 0e12e445-3aea-4195-bc41-e52d725cb488 · inbound

Quantifying Memorization Across Neural Language Models cites this paper.

Quantifying Memorization Across Neural Language Models Extracting Training Data from Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:04:59.731563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:04:59.678438Z digest=sha256:89046e2cbdceb59bfbcec9fed4d510991a93fdf14e54f1e409120c289005b87b

Observation 78f2a4f4-10c0-4085-a8cd-f4364edeca01 · inbound

Scaling Laws and Interpretability of Learning from Repeated Data cites this paper.

Scaling Laws and Interpretability of Learning from Repeated Data Extracting Training Data from Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:52:40.461771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T15:52:40.335080Z digest=sha256:2d6a518fadcf3b6e950c515d531d931448ba4659780d379e7c112afebfb56443

Observation 862123b6-3250-4082-8f30-c978da3b3fc5 · inbound

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cites this paper.

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned Extracting Training Data from Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:38:08.468054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:38:08.362920Z digest=sha256:8970f07e3fe1d1b1f6f60971847e338cd7b85bc04bac777fe3ea890587ed32df

Observation 54ae7179-ff65-4628-80c9-4d5bd6586ed9 · inbound

MusicLM: Generating Music From Text cites this paper.

MusicLM: Generating Music From Text Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:27:06.256783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T14:27:05.198402Z digest=sha256:9e15edc612bbd16fd164b1ce6ed6740cf917f1265e8068e5624e0a877931a255

Observation ae0ef567-9f81-41fc-a125-e94baeaec035 · inbound

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions cites this paper.

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions Extracting Training Data from Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:23:52.895916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:21:49.775278Z digest=sha256:cb6d69dbba97e3f78bf7cec8378b5478c3bdb076f217e3221d64ff17793a5acc

Observation 8fd242df-c7e3-41f4-83a3-3354df4306fa · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Extracting Training Data from Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.740850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:c68f84dd9ec6a2129e3c6126c30441d7c23245c2d0498ba70d70b6feb364946a

Observation 07b45183-d87e-42d3-aee1-ba78389b23f1 · inbound

Towards the Anonymization of the Language Modeling cites this paper.

Towards the Anonymization of the Language Modeling Extracting Training Data from Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:32:39.434987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T06:28:16.975305Z digest=sha256:a65cfbc6f8d835ccf38a9463c910c643dc26b5123f70457e30f93f4c29db21f2

Observation 62fcb607-fd74-4d3c-ad40-67c0d7f50894 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Extracting Training Data from Large Language Models

Reference 253

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:11:37.303275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:37923eca85dd87adef4540b9a5875e308d3ea3b285ecc44d21855b046c082f87

Observation 73256e9c-8ad7-463d-8010-e673d7e5b027 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning Extracting Training Data from Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:25.072025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:25.072025Z digest=sha256:aa849bedc54e735d94645e056a3206b47b8ab9f9ee259ad4a8dfb5878cec5d24

Observation 65626d80-068e-4c69-a336-5968e501f862 · inbound

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation cites this paper.

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation Extracting Training Data from Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:34:30.348594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:34:21.657999Z digest=sha256:8294b6e00b8b378e56c4a511a79682322438b0c70975de762e1f17f2f0a6a1e6

Observation 63487873-dc3e-4722-88fe-87cac685c658 · inbound

Approximating Language Model Training Data from Weights cites this paper.

Approximating Language Model Training Data from Weights Extracting Training Data from Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:58.403904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:58:58.403904Z digest=sha256:66e120d7b056ad4f65c4794c1766f80f1da7cb5d098bec3af260be502dcb9a03

Observation 337899db-b046-4a2b-9411-196bf7d67a2e · inbound

From Teacher to Student: Tracking Memorization Through Model Distillation cites this paper.

From Teacher to Student: Tracking Memorization Through Model Distillation Extracting Training Data from Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:10.559721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:10.559721Z digest=sha256:5cf499746efc0a06c3acd22f958d46f33647b24aa806fb28f9c2bfd022b05a32

Observation 42bf4c06-3c92-4a00-9aa0-b9d9be79d119 · inbound

Low-Perplexity LLM-Generated Sequences and Where To Find Them cites this paper.

Low-Perplexity LLM-Generated Sequences and Where To Find Them Extracting Training Data from Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:58.634485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:44:58.634485Z digest=sha256:1724fe894b73398b3f5ef5ae22be136c4f0df668363d3b806abbd41c8d83078b

Observation 6ea55d67-c336-4e53-831c-47d55a3e9ae5 · inbound

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models cites this paper.

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models Extracting Training Data from Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:55.141023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:55.141023Z digest=sha256:3352c632dc7c8b0cddc2e92915b77ab4c8a2cca9b771a2ef160044499975dbdf

Observation e941e346-f608-4269-aa6c-193a4c956cbc · inbound

Adversarial Machine Learning Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack cites this paper.

Adversarial Machine Learning Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack Extracting Training Data from Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:00.983161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:31:00.983161Z digest=sha256:81322fdf12301ea1248f268096139553aab1a3c0ab31f34b316c970a310ac1c5

Observation 77948df8-f2ae-4a6a-9eeb-3f0a4ab7d5c2 · inbound

On the Performance of Differentially Private Optimization with Heavy-Tail Class Imbalance cites this paper.

On the Performance of Differentially Private Optimization with Heavy-Tail Class Imbalance Extracting Training Data from Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:00.426132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:00.426132Z digest=sha256:0bcaacabc720a1e83fa1ee0ab63c3385a824393d717578d5e132ad332b4a8a6a

Observation aa9c0ed5-5f41-4d48-941c-dccbe7fe0fdc · inbound

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI cites this paper.

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI Extracting Training Data from Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:55.618154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:55.618154Z digest=sha256:af29ec06583ceb9d176c1aea7d862df52e10e5ff221366522ffefe2343a2a4e0

Observation 45843fde-c633-4720-a317-5c86e2576e11 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Extracting Training Data from Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.822425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.822425Z digest=sha256:969a0fadfe3749d01d4de34356b5f7a9b153ccc3a1063fc4d954dd1d0558f3fd

Observation cd504875-9749-41af-a9f8-8cfebb784e8d · inbound

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage cites this paper.

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage Extracting Training Data from Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:50:19.605008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:50:19.605008Z digest=sha256:f0685040d4b2ae4168870c455f971d9bb103db08eb7d8b7382bd18bee4be1e24

Observation 5a4074d2-053f-4fce-b7a7-fa03fb5fe38e · inbound

Evaluating Differentially Private Generation of Domain-Specific Text cites this paper.

Evaluating Differentially Private Generation of Domain-Specific Text Extracting Training Data from Large Language Models

Reference 2650

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:03.247046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:09:03.247046Z digest=sha256:dfe89199272c8e144df22a87688f13bd41727227c1f7cd8a3a4baae6a4d21cb6

Observation 178f3650-d0a3-4e33-8ccd-087eef77de7b · inbound

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants cites this paper.

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants Extracting Training Data from Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:46:41.159734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:46:41.159734Z digest=sha256:206456f9836d14463614c79cc78988f8071685fcae29786b884601df487d2bca

Observation ffe1c536-0a58-4f78-b42e-25b0318b13be · inbound

SynBench: A Benchmark for Differentially Private Text Generation cites this paper.

SynBench: A Benchmark for Differentially Private Text Generation Extracting Training Data from Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:46:37.468783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T16:46:22.479895Z digest=sha256:1938a0df7a9c7ae2c406c5a93153468fa9b6e468f6f55186f0ea7656b459650d

Observation 1b70cc15-f9fc-40f1-864b-0fc5c61e68b4 · inbound

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation cites this paper.

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation Extracting Training Data from Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.851547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:47:31.667427Z digest=sha256:e92afc8eb95e8df2ef4aa34749cda188fcee9aad207e0bd0b4b2351cc4ddf3f2

Observation a00801a1-c0c0-4373-9b48-aa540bf3dae9 · inbound

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts cites this paper.

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts Extracting Training Data from Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.805002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:44:26.931344Z digest=sha256:cf425db2f20d0b3f88019c593f527658899a6d680d1e2bc324fb767f6fda6f2b

Observation 6dc63095-16dc-4a9c-bda3-7ae3027a0fea · inbound

Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies cites this paper.

Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies Extracting Training Data from Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:34:07.373969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:26:34.487400Z digest=sha256:66cc2d01a84d02108b901a50307bab778badc86edbb82823f97393644ea8c2f1

Observation 5b4fee91-8675-4ad1-9c0a-12203b86194f · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Extracting Training Data from Large Language Models

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.389604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:e70877b69a3425b08f662ba2c355d989744928555f2bbebedd4c8e22b52b2c75

Observation 89517f86-bac3-429d-b231-14029185edba · inbound

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model cites this paper.

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model Extracting Training Data from Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:17.879632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:12:20.534893Z digest=sha256:ff0af8da14b1195d6036f9c117b835b9f1bdb12563a2eb2c53de6991f8e4b07d

Observation 4f74794c-9c4d-4e90-b0d3-ef5718da6854 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Extracting Training Data from Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.288375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:4858a7a6c1fc76e43a4b27ede5ac01b2694b4621c588b558d7387ac5a9c6d842

Observation 57b34014-a152-497f-914f-558789c1199b · inbound

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems cites this paper.

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems Extracting Training Data from Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:01:05.960115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T04:58:37.780195Z digest=sha256:df5f901f4bdf69999610487f2423b297da8c616eeb9cf3b9efb32a7f5a78f2c1

Observation f9fbb4f0-4061-4a72-9d45-69a78e5d0c7e · inbound

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing cites this paper.

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.664060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T22:37:02.708284Z digest=sha256:ebf330aeb998fcdd34f7ca041f6f31a388ad8d5fae894893a802c037b1d08446

Observation 40c45581-7f6a-4603-a973-ad6231aaf9b3 · inbound

MRMMIA: Membership Inference Attacks on Memory in Chat Agents cites this paper.

MRMMIA: Membership Inference Attacks on Memory in Chat Agents Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:27.072292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:04:53.511344Z digest=sha256:81f18040f1c5ccf2e0120f0c4b30990e881ff7f16c018fb1b76b5ada89350494

Observation d64b9315-cf10-4cbb-b788-e426d461e0a9 · inbound

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing cites this paper.

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:56:23.234607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T13:49:47.542697Z digest=sha256:071164e020172deec5ec6bea78fed8636429743ac5e757135c48a3aa6cd3535d

Observation 1468ffd0-0d5b-4fbc-8846-251fc95b7cbf · inbound

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails cites this paper.

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.057543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:06:31.341519Z digest=sha256:47857e63bf2eb69a2bdc74aafac01346e2e06fbd7d910a76c33c20f9d4c34663

Observation 80ad5aa8-0703-40ed-8ccb-81d3bd7d3de1 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Extracting Training Data from Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.468052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:5d160f6a9148a57286302b1927da89bcd37691b8e1a06c15a0c04714b98226d2

Observation f30c694f-27ce-46d4-9fc0-2cd79110ae11 · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Extracting Training Data from Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.357949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:33ee5359c50fda83ab590ebfa5296b92c6a034329d4d91d276e7efc7784edeec

Observation b6dac66e-c769-457e-8edf-e64ccde2034f · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Extracting Training Data from Large Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:07:30.231259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:75736c6cc813ff7b57b71891c94dfbe9b85cee89ec1d2ad4979184a2f93873cc

Observation e2d36494-b0a8-4b0d-8507-d9c8ae0a36cb · inbound

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans cites this paper.

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans Extracting Training Data from Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:36.687299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T09:00:23.558245Z digest=sha256:0aa677fa21c729358bbc0c2dc0e59c483b21a76aa38cd9fbb8a61ef9d76b61da

Observation 48cfd250-1281-4f89-a3b2-72bd4189fdcf · inbound

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents cites this paper.

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents Extracting Training Data from Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:58:06.926713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:11:42.204778Z digest=sha256:3ecd2cbc55395032aeff16391a3e8d80805ae334534386ed24b16d5a350a79c1

Observation fb4f126f-1954-4a98-80d7-4e31d24dede2 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:47.182482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:e2487924076105a11708333b52f4a7149b7cf4a115d8bb9fab5d2294c7a522cd

Observation 99fef884-5b0d-4d97-9e02-13ed9e2c4cf6 · inbound

Exposing the Illusion of Erasure in Knowledge Editing for LLMs cites this paper.

Exposing the Illusion of Erasure in Knowledge Editing for LLMs Extracting Training Data from Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.276219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:10:39.422141Z digest=sha256:54c8eafec345afbc75a7724505692ead54c111ebb0e42e44fcdea463007a7526

Observation ec1a872e-ec30-4beb-b1a3-9eb6d7dd83c7 · inbound

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents cites this paper.

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents Extracting Training Data from Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:09:53.281660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T04:29:16.386339Z digest=sha256:2211d4c1a2f5a719b0eb0b9ec21bfab5020df27fc16e38b1c228cc2bf9a9b70b

Observation ffdbb1a8-433a-472a-900d-5d4097d0fa7d · inbound

AI Native Games: A Survey and Roadmap cites this paper.

AI Native Games: A Survey and Roadmap Extracting Training Data from Large Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:58.758903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T12:59:20.908659Z digest=sha256:a6068585106644a0800c2e474974211d66cc7fa59908ddb47238a66725fd53b9

Observation 4d8c8d02-19f8-4a42-88a7-cde6e538eab4 · inbound

AI Native Games: A Survey and Roadmap cites this paper.

AI Native Games: A Survey and Roadmap Extracting Training Data from Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:07.729486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:07.729486Z digest=sha256:f0035c4ab8141986f2dfb2ab0e24eaba1e9b2cc255b25d734e7ab9ead58db7d4

Observation 41d6405e-1aa9-43f1-8a6a-e244df868e0e · inbound

Auditing Forgetting in Limited Memory Language Models cites this paper.

Auditing Forgetting in Limited Memory Language Models Extracting Training Data from Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.312170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T13:09:34.491580Z digest=sha256:858ed07b55bcc1ba2fde356134c4fbfa8017e86cf806e6cdc47abfa25d3e890a

Observation 724b7aaa-782d-4928-b56a-e74cdf712e32 · inbound

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification cites this paper.

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification Extracting Training Data from Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:11:43.482286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:11:43.482286Z digest=sha256:0e3d1442fada4fd3c278a3a1c6bf02ee5a9c699b7a62bb9dfa0315c21088509f

Observation 9c1c3df4-b145-41bb-a08f-62af760ed21c · inbound

DECAF: De-Clustering for Adaptive Representational Unlearning cites this paper.

DECAF: De-Clustering for Adaptive Representational Unlearning Extracting Training Data from Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:33:49.808882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:33:49.808882Z digest=sha256:21f9f90ea1f7f5dc85a5ec6d5f630e98de99759ad94330ef8bf5e493cc0db2a7

Observation c08c5f00-e512-4bec-91e8-c1bbee4adaf8 · inbound

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization cites this paper.

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization Extracting Training Data from Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T02:28:36.726398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:28:36.726398Z digest=sha256:4b17e8d5e24c51fc8747d52bafdc5f8c219c099e0ac016ccf618a68fe4e19067