Pith. sign in

Paper Citation Record · LEDGER

Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2404.02936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.02936 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:11:59.861601Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b20cd380-7302-4846-bcfb-40d6757d1a3e · inbound

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models cites this paper.

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T04:11:59.861601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:11:59.861601Z digest=sha256:b1c6fc5fd62baa95832dbf13ef4fa71f642d06399bd13bb245866ce7f115d23c

Observation 5207c1a4-dda9-4edf-9908-bf130e0cc506 · inbound

DUSK: Do Not Unlearn Shared Knowledge cites this paper.

DUSK: Do Not Unlearn Shared Knowledge Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:25.147132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:25.147132Z digest=sha256:5a17623c847b3b8790317282549a72ca3b956231c5c0c72440156cd54add0738

Observation 548de2d1-4350-4c21-b212-3e1c67f612a8 · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:06.177607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:06.177607Z digest=sha256:080349286fd02b4bc423cbd5b136c12e01590ab2ddfb97a144106e132d611b7c

Observation 1d64ba19-d40c-4db7-b05e-9791aa2bd424 · inbound

Membership Inference Attacks on Sequence Models cites this paper.

Membership Inference Attacks on Sequence Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.967563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.967563Z digest=sha256:148077bfb0d162d5348d691ef15519d4accba4d0cb08fb2d4c48ccf6030302c6

Observation 54e7395b-2eba-4243-8bec-cb458cf5e60b · inbound

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems cites this paper.

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:16.451624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:16.451624Z digest=sha256:95e48e19450da3f8723dead9a25ffe684087d2d4ee737d616dbdf913867cfd16

Observation 8c09842f-e619-420a-9d09-bbaf013e431f · inbound

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks cites this paper.

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:17.063199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:17.063199Z digest=sha256:9407d92d870ff6ba131b076c60f73afc8cf4410e44ec1f32bb026fab5e98dac8

Observation 27c22f6c-8553-4cf2-9002-a8896860bd67 · inbound

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework cites this paper.

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:15:21.428831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:15:21.428831Z digest=sha256:b92304e4ef211e7b19a0ff74b6db146a02ef4b3b79af65742e864c7ff6c1a62c

Observation 34c00687-a372-46be-8ae7-bcb5347cb0e3 · inbound

Investigating Training Data Detection in AI Coders cites this paper.

Investigating Training Data Detection in AI Coders Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.839169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.839169Z digest=sha256:ea102409f6b50aa446b379b73fc17177109511c687f38680f230d8ef2a5cce5a

Observation 034fc69a-6222-4535-ab99-2cf42a1c8c0d · inbound

Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis cites this paper.

Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:29:27.157115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:29:27.157115Z digest=sha256:cf60711196512d41be8c9fc87d28196a2e30906f54dfa82d7206c42f3565661f

Observation 8052ae0b-864b-4d1b-a8ad-db208d0d0cdc · inbound

Membership Inference Attacks on Tokenizers of Large Language Models cites this paper.

Membership Inference Attacks on Tokenizers of Large Language Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:22.454938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:22.454938Z digest=sha256:6417d5f4f2a38e434bbf617d0b778869e89466406d4984864eeabe349767a5fb

Observation 36584a79-e6c1-4784-b819-5ccbfe15c863 · inbound

LLM generation novelty through the lens of semantic similarity cites this paper.

LLM generation novelty through the lens of semantic similarity Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.311136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.311136Z digest=sha256:559f9d1d89ca932a2ea5c6c23af19c06415cd2b6ee6f3dbe7d673948bf844fe0

Observation 625fbecc-31a9-4d15-8ad8-89a18e076274 · inbound

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards cites this paper.

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:40:17.766089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:37:54.702010Z digest=sha256:2bd6152af9e6fb1b6ce9b946debefd9cd911224d47bd0acf506a1694be91418e

Observation ceca5c65-e301-4a5f-a529-51ac3e4385ff · inbound

Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models cites this paper.

Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:44:50.219620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T11:42:18.076004Z digest=sha256:72ff6b701badba8f31dc1583efe6bae844e74312eceecc246baf5cc67e1290fe

Observation a7adb287-207d-4551-b810-bc240ae08fef · inbound

Learning the Signature of Memorization in Autoregressive Language Models cites this paper.

Learning the Signature of Memorization in Autoregressive Language Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.478145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:53:10.396785Z digest=sha256:67e8d0d3919b5d509bd21a1d7ec0fc3f030787e0077afcf025db0308b3cad889

Observation df7fb1b4-698c-4156-b89d-f663c5452565 · inbound

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs cites this paper.

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.938818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T15:32:05.976335Z digest=sha256:d4aa16f40d99c4b914be70d490eddfcc92510ee7a492642a35923a2c9e5687f1

Observation 00cc933f-07da-414e-b0a8-d626ea453f1b · inbound

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning cites this paper.

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:56:27.596710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T11:27:02.902720Z digest=sha256:09bac3214dc9a51e1c5ecc38fca85d701193f9c5c2e7ce8f32242129362b202e

Observation e07449e6-5615-4e5d-830a-d4a80a3db36c · inbound

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning cites this paper.

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T12:32:56.389581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:32:56.389581Z digest=sha256:573dcbc768af77be2370f99adc0fd834ad614393c21f2033a7be7d136936c7f0

Observation a0c2d51d-8867-4505-9a6a-2d199d838521 · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.329479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:bdf14994378a233ec08cea7c56fc5a570da7ddfa5e85c08d3b542030883d6076

Observation f02e6a18-0bbf-4e9c-b300-0dbdb7bc3c4f · inbound

Amplifying Membership Signal Through Chained Regeneration cites this paper.

Amplifying Membership Signal Through Chained Regeneration Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:25:26.760412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:22:08.140403Z digest=sha256:7acf3a1a7dcc97ef5ccec06f0c16700304408d227b51ee5e7d60b2a1e88b1751

Observation 3d34a930-1de0-4218-a169-0891cc3ef470 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 188

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:37.193909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:37.193909Z digest=sha256:f436ef77c09f5b76ff6a69e91a25f874c6dc5f608aff54d21db052de3aa2b112

Observation 3c0af4f0-f221-4eca-aa36-a9b188f75ecc · inbound

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration cites this paper.

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:35:50.513439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:35:50.513439Z digest=sha256:72f2996cb7b4d0b45bea438d9c9029fc49bb6c144f7b535b42b56c64dc278625