Pith. sign in

Paper Citation Record · LEDGER

Membership Inference Attacks on Tokenizers of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 1 inbound Pith citation observation for arXiv:2510.05699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.05699 v4

Coverage vector

measured 100 of 110 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:21.490138Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T14:12:14.160789Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T14:15:54.917491Z

Reference resolution

100 of 110 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89ca683b-9c28-46d2-b144-17a72af8f5ff · outbound

This paper cites Deep learning with differential privacy.

Membership Inference Attacks on Tokenizers of Large Language Models Deep learning with differential privacy

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:08.690312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:08.690312Z digest=sha256:8edb28217562f60d6bf7e13695e4ee599f0c6449e1f6ea23b7f26058deeb064f

Observation 3cc7c33f-5f83-47a8-96a9-6f0b7ca7df56 · outbound

This paper cites An information-theoretic perspective of tf–idf measures.Information Processing & Manage- ment, 39(1):45–65, 2003.

Membership Inference Attacks on Tokenizers of Large Language Models An information-theoretic perspective of tf–idf measures.Information Processing & Manage- ment, 39(1):45–65, 2003

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:08.782017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:08.782017Z digest=sha256:99c7c24223209662b5678f6c4b23408223af8a2d85443127d8dfd24edda12c6a

Observation 6ec36006-0b10-487b-a845-c462faa331b5 · outbound

This paper cites Judge allows new york times copyright lawsuit to go forward.

Membership Inference Attacks on Tokenizers of Large Language Models Judge allows new york times copyright lawsuit to go forward

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:08.943368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:08.943368Z digest=sha256:9d4c20957db245990afa9e90327b40be7c57749ebf574b93abd602b14c4fb6f1

Observation 10c4a399-da1b-4a44-b4dc-e08208547665 · outbound

This paper cites Tokenizer for anthropic large language models.

Membership Inference Attacks on Tokenizers of Large Language Models Tokenizer for anthropic large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.100331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.100331Z digest=sha256:eceee19a5673a1061a63931d77b341fc113f5ca3a943a074c7756bd120c08b69

Observation c03111da-bd68-4471-b239-a205630f5cda · outbound

This paper cites Claude opus 4 & claude sonnet 4.

Membership Inference Attacks on Tokenizers of Large Language Models Claude opus 4 & claude sonnet 4

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.258490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.258490Z digest=sha256:f33fca631c4cd303b00dc5a831b25dcbfe09e199de82782937b4981b459c2025

Observation 13c85151-3bc7-475a-b9ec-23eb93fc1a04 · outbound

This paper cites An efficient recommendation generation us- ing relevant jaccard similarity.Information Sciences, 483:53–64, 2019.

Membership Inference Attacks on Tokenizers of Large Language Models An efficient recommendation generation us- ing relevant jaccard similarity.Information Sciences, 483:53–64, 2019

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.382455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.382455Z digest=sha256:5270e94266681df93c197c5043aa14725922274763d95db5e9e8bde90956cafc

Observation 9e6aa735-4fed-4836-b236-4946d666d47e · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Membership Inference Attacks on Tokenizers of Large Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.498777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.498777Z digest=sha256:89ba771f94b7b69c9a0d46d45b10df7f7de27f94592a8fa91b53ff61367cfb29

Observation a65b57f8-2800-4351-bfa6-d8649e7593aa · outbound

This paper cites Pythia: A suite for analyzing large language models across train- ing and scaling.

Membership Inference Attacks on Tokenizers of Large Language Models Pythia: A suite for analyzing large language models across train- ing and scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.592846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.592846Z digest=sha256:fe170001f28d06be6537e3199885e269ad41fbf2c37ba6492314f4ad38177c26

Observation 414d1b89-c345-4fd5-af62-d423de99ebe4 · outbound

This paper cites Gpt- neox-20b: An open-source autoregressive language model.Challenges & Perspectives in Creating Large Language Models, page 95, 2022.

Membership Inference Attacks on Tokenizers of Large Language Models Gpt- neox-20b: An open-source autoregressive language model.Challenges & Perspectives in Creating Large Language Models, page 95, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.743250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.743250Z digest=sha256:5a5f70d61f0dd1f82ec48b52bf08d372fb2d382e60041cb27bc16e4c7c7e819f

Observation 2f6d474c-fcd8-4d0f-9c32-015dcfb26820 · outbound

This paper cites Language models are few-shot learners.

Membership Inference Attacks on Tokenizers of Large Language Models Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:09.913728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:09.913728Z digest=sha256:befba657751c337ecb4467f8407d5a26549d3d35c7c582c7070e2d8ce4ab6800

Observation 098d1c8e-010d-4e15-b630-b9943137dbcf · outbound

This paper cites Member- ship inference attacks from first principles.

Membership Inference Attacks on Tokenizers of Large Language Models Member- ship inference attacks from first principles

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.078307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.078307Z digest=sha256:5d8c2f3b232f704072599fc7dcaf66014d7dd94cb538f822e61755d5e72a4cca

Observation 803170ef-c58f-4191-a968-bc6846138641 · outbound

This paper cites Extracting training data from large lan- guage models.

Membership Inference Attacks on Tokenizers of Large Language Models Extracting training data from large lan- guage models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.298498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.298498Z digest=sha256:909ad5587efa6821f3c1bcb7f20d7958d50c006801a5421557b32593f0e7c16d

Observation 12b3022d-df34-4093-bfe0-f84599291d54 · outbound

This paper cites an unresolved cited work.

Membership Inference Attacks on Tokenizers of Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.445591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.445591Z digest=sha256:cf2b6ae65a8f34e8333dadbdb3f8dfee3128988f87ce8028a88195a992ceb2ff

Observation ec55dba2-a01f-4865-bf65-1683e959af23 · outbound

This paper cites Evaluating the Dynamics of Membership Privacy in Deep Learning.

Membership Inference Attacks on Tokenizers of Large Language Models Evaluating the Dynamics of Membership Privacy in Deep Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.478201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.478201Z digest=sha256:7030b0cf107cd2576c15a655ba8613bf49ce88f082872c14180ad671df87e33d

Observation 9234e720-7ce7-4196-ad70-afc777453ef6 · outbound

This paper cites How contaminated is your benchmark? measuring dataset leakage in large language models with kernel divergence.

Membership Inference Attacks on Tokenizers of Large Language Models How contaminated is your benchmark? measuring dataset leakage in large language models with kernel divergence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.540157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.540157Z digest=sha256:c461ddf620b9528ca2d4398153a93031b9ced7cc5266cb87cea111f89a053703

Observation 6149220c-deb4-4710-b6dd-8384bb28e7d2 · outbound

This paper cites Label-only membership inference attacks.

Membership Inference Attacks on Tokenizers of Large Language Models Label-only membership inference attacks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.699085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.699085Z digest=sha256:b20dfdd3688b473a812cea077db49aeaa93849f9182304dca7c8df281a79dcec

Observation 4eebd633-09a4-4821-a45d-6414e311e49b · outbound

This paper cites Power-law distributions in empirical data.

Membership Inference Attacks on Tokenizers of Large Language Models Power-law distributions in empirical data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:10.826580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:10.826580Z digest=sha256:1dee7249a861606d98e9c7a836741fd5d7150fd4477cf8da46b586f5ab05ce48

Observation e0a29057-cc14-4b30-995e-1b10e1a08329 · outbound

This paper cites an unresolved cited work.

Membership Inference Attacks on Tokenizers of Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:11.021457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:11.021457Z digest=sha256:0743cfed9cbe303ea1461ae5fe149d433600209e3454c67104f0eeda4283dacb

Observation c1d21558-101f-4025-a1e0-a8ebbba5cea8 · outbound

This paper cites Getting the most out of your tokenizer for pre- training and domain adaptation.

Membership Inference Attacks on Tokenizers of Large Language Models Getting the most out of your tokenizer for pre- training and domain adaptation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:11.293865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:11.293865Z digest=sha256:a2be36995f36ca05447ccee2378ac3117d821295bfefa8f4d0ce0a963aac0e78

Observation 4bffe1b0-a320-48f4-a8f1-836002e189f3 · outbound

This paper cites Blind baselines beat membership inference attacks for foun- dation models.

Membership Inference Attacks on Tokenizers of Large Language Models Blind baselines beat membership inference attacks for foun- dation models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:11.478113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:11.478113Z digest=sha256:36628fc2069e640f2cdde6f37dafa073d73d30da690b8f55ca4ac838f2d944d9

Observation c4d2d478-1ac0-4395-a50d-f899d4dec326 · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

Membership Inference Attacks on Tokenizers of Large Language Models Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:11.686039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:11.686039Z digest=sha256:607ec2c5efb15183297211a2af7a5b07ba8cfab9c5a8bb5d6c6ec7920b3184d8

Observation cb1496d5-0a47-4cfe-86f0-5cad142922ec · outbound

This paper cites Dp-forward: Fine-tuning and inference on language models with differential privacy in forward pass.

Membership Inference Attacks on Tokenizers of Large Language Models Dp-forward: Fine-tuning and inference on language models with differential privacy in forward pass

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:11.813613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:11.813613Z digest=sha256:94b07282742f794b5a6c0aeba1386c30cb8ea8774a0779060aeb976c24548579

Observation 30892f90-a83a-4d54-843e-b54ff8c7d596 · outbound

This paper cites Cascading and Proxy Membership Infer- ence Attacks.

Membership Inference Attacks on Tokenizers of Large Language Models Cascading and Proxy Membership Infer- ence Attacks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.002596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.002596Z digest=sha256:c2cf23ac2528aafc821ce05ecc32b3a81162daf6c6d6f061aec9c2fab8ad7e7e

Observation 8fc3d85b-3406-48db-8572-b56659f47b96 · outbound

This paper cites Systematic Assessment of Tabular Data Synthesis.

Membership Inference Attacks on Tokenizers of Large Language Models Systematic Assessment of Tabular Data Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.207549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.207549Z digest=sha256:521ed3bb637b4913962b390961dfba41d0bf603f48b89f45875338a9e67d1507

Observation 685a4178-e04e-4c8f-9717-d50a35ac99ff · outbound

This paper cites Do membership inference attacks work on large language models? InFirst Conference on Lan- guage Modeling, 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Do membership inference attacks work on large language models? InFirst Conference on Lan- guage Modeling, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.416804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.416804Z digest=sha256:8b11f6176736c55bdc4f657308d99cac36e4a9f8ddfa23ac8a04884699054faf

Observation 720d92b9-6782-4471-8692-a3be8a943d0e · outbound

This paper cites De-cop: detecting copyrighted content in language models training data.

Membership Inference Attacks on Tokenizers of Large Language Models De-cop: detecting copyrighted content in language models training data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.556633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.556633Z digest=sha256:ef8617d33534f2829c19f5ee1028022c9e99548400b0cbc765941d4c6cbdb245

Observation 227687d3-f9dd-40d6-bdb9-a6746c42256e · outbound

This paper cites Differential privacy.

Membership Inference Attacks on Tokenizers of Large Language Models Differential privacy

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.670935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.670935Z digest=sha256:f7b1d3fb0b45000220c7aed64914ebabcec4d38a7398088207dd0618b2583bbc

Observation 67d56b7d-f175-4eee-9a24-7cfb1ef9743f · outbound

This paper cites Analysis of sparse bayesian learning.Advances in neural information processing systems, 14, 2001.

Membership Inference Attacks on Tokenizers of Large Language Models Analysis of sparse bayesian learning.Advances in neural information processing systems, 14, 2001

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.789453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.789453Z digest=sha256:822d13017feb6fc99a6be95d5041a4e055e718bb71b3bb3bda2242475b68294c

Observation d2e8c732-49f8-4da3-9e9f-3c812b3b3b0f · outbound

This paper cites Privacy in pharmacogenetics: An {End-to-End} case study of personalized warfarin dosing.

Membership Inference Attacks on Tokenizers of Large Language Models Privacy in pharmacogenetics: An {End-to-End} case study of personalized warfarin dosing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.883547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.883547Z digest=sha256:2397d95b67df12c1a1993e5b4e5479e92484ac9a0b931dc9ba4f54fe5010b149

Observation 780ceb07-6693-4b17-9a11-b34a1e5b28bd · outbound

This paper cites Label inference attacks against vertical federated learning.

Membership Inference Attacks on Tokenizers of Large Language Models Label inference attacks against vertical federated learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:12.966221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:12.966221Z digest=sha256:e0a739d5924a17a5d5f9da5ebf755f1df17087b382f09b5bcb5fd5b356420a64

Observation 092847cd-3848-48da-80c5-910353f0d4cc · outbound

This paper cites Zipf’s law and the growth of cities.

Membership Inference Attacks on Tokenizers of Large Language Models Zipf’s law and the growth of cities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.066661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.066661Z digest=sha256:27aa32adeaa5cc36d401be2d81403122f2060487f474d907a143a162c85e9eb1

Observation f3b2c762-26a8-472b-a9f1-1f29ca244880 · outbound

This paper cites Investigating the effectiveness of bpe: The power of shorter sequences.

Membership Inference Attacks on Tokenizers of Large Language Models Investigating the effectiveness of bpe: The power of shorter sequences

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.176008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.176008Z digest=sha256:d3f4bca860dc3afe3c8364dd4af66758d30973c03307e5aad0dd99e9fa7664ed

Observation faae467b-149c-4f95-a9fd-77ea20f46e57 · outbound

This paper cites Counting gemini text tokens locally with the vertex ai sdk, July 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Counting gemini text tokens locally with the vertex ai sdk, July 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.258531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.258531Z digest=sha256:745224c286eed2a2123aaaabd23c660393c826f9c6e5cfdd9c35444ba53b2030

Observation f8f08380-5760-45e7-beb1-de4c609416bc · outbound

This paper cites Likelihood-based diffusion language models.Ad- vances in Neural Information Processing Systems, 36:16693–16715, 2023.

Membership Inference Attacks on Tokenizers of Large Language Models Likelihood-based diffusion language models.Ad- vances in Neural Information Processing Systems, 36:16693–16715, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.383669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.383669Z digest=sha256:87ad4d4d0dcc4d3044934bce9cff52a16c53e23a5c5a8fa3e26d53d94acddb1a

Observation 67ad8c75-cd09-42f4-ac54-8e01f62f6e3a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Membership Inference Attacks on Tokenizers of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.479968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.479968Z digest=sha256:03a98d3569156004e9b48894488ab91de1db4a268cf4fa08feb1aa174b91ddd5

Observation 6e9ced76-0fb6-4eb2-ad8c-84b82d1e6ba1 · outbound

This paper cites Weird gpt-4 behavior for “davidjl”.

Membership Inference Attacks on Tokenizers of Large Language Models Weird gpt-4 behavior for “davidjl”

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.553514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.553514Z digest=sha256:08b653560cce16561b0a40dccc5c483924d10d65c3e200645a227be019f1cad3

Observation 7ac65f64-94ce-4f18-9d6f-64dcc68ccc30 · outbound

This paper cites Data mixture inference attack: Bpe tokenizers reveal training data compositions.Advances in Neural Information Processing Systems, 37:8956– 8983, 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Data mixture inference attack: Bpe tokenizers reveal training data compositions.Advances in Neural Information Processing Systems, 37:8956– 8983, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.694787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.694787Z digest=sha256:4989a0f08abaca4871e8c0e782614618e69bb277eeb795b1715e7d279c60609f

Observation 22c71755-5ae2-442b-8be2-015103f505d2 · outbound

This paper cites Data mixture inference attack: Bpe tokenizers reveal training data compositions.

Membership Inference Attacks on Tokenizers of Large Language Models Data mixture inference attack: Bpe tokenizers reveal training data compositions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:13.859360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:13.859360Z digest=sha256:6a9c4d866129b290108b95809f2c4252be9aeea02a37755911f1406370fb91e9

Observation a866abed-7977-441a-92de-32ce83a00e20 · outbound

This paper cites Strong membership inference attacks on massive datasets and (moderately) large language models.arXiv preprint arXiv:2505.18773, 2025.

Membership Inference Attacks on Tokenizers of Large Language Models Strong membership inference attacks on massive datasets and (moderately) large language models.arXiv preprint arXiv:2505.18773, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.044399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.044399Z digest=sha256:06a9eb002e5e2dd3205d0ca6beab0e0f4c6c5ee2c852284abce1aa16985ce34e

Observation e39349ed-285e-4f0b-a714-160e52d70c93 · outbound

This paper cites Deep residual learning for image recognition.

Membership Inference Attacks on Tokenizers of Large Language Models Deep residual learning for image recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.196443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.196443Z digest=sha256:fc2e8b21f49a0992e7eb1682df75ca0e836fcf940a16504ab3e66445318e74f1

Observation 7a66f518-0e33-450b-8fcd-5e29b7c89e84 · outbound

This paper cites To- wards label-only membership inference attack against pre-trained large language models.

Membership Inference Attacks on Tokenizers of Large Language Models To- wards label-only membership inference attack against pre-trained large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.338840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.338840Z digest=sha256:c940703f2c0cb74b558b44f26e5da36ffa9ab289984079104608391e96fb4fdd

Observation da5416fe-9973-4fc6-b64e-e5260c2393a5 · outbound

This paper cites Membership inference attacks against vision-language models.

Membership Inference Attacks on Tokenizers of Large Language Models Membership inference attacks against vision-language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.477396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.477396Z digest=sha256:24abbd561d3a12b4a913ada16b8bfad2bfff0882181ea9e667cdaababe8b82b9

Observation db2e21bf-a47c-40d9-8aa7-39213b4fe42f · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Membership Inference Attacks on Tokenizers of Large Language Models Scaling Laws for Autoregressive Generative Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.629106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.629106Z digest=sha256:5ec3b8a1d4d342794fba0fbec0479e2cfab69d33bab92222785592076a26ab23

Observation 42950e4d-d92d-4c75-b40f-e824e7696a2c · outbound

This paper cites Training Compute-Optimal Large Language Models.

Membership Inference Attacks on Tokenizers of Large Language Models Training Compute-Optimal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.813967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.813967Z digest=sha256:3c3cdb517c68c12852177e00e6639dc274a46eb3512ee54d6945c5cb48e001c8

Observation 0ae86809-87e4-450e-94bf-11e5cae853c7 · outbound

This paper cites Damia: Leveraging domain adaptation as a defense against membership inference attacks.IEEE Transactions on Dependable and Secure Computing, 19(5):3183–3199, 2021.

Membership Inference Attacks on Tokenizers of Large Language Models Damia: Leveraging domain adaptation as a defense against membership inference attacks.IEEE Transactions on Dependable and Secure Computing, 19(5):3183–3199, 2021

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:14.917804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:14.917804Z digest=sha256:786144d86e0f67daaa341c89d9c64749449d636eebda80d33a2349580e0f5e68

Observation 2c985725-1bba-4182-8cab-403ad9c905a1 · outbound

This paper cites Over-tokenized trans- former: V ocabulary is generally worth scaling.

Membership Inference Attacks on Tokenizers of Large Language Models Over-tokenized trans- former: V ocabulary is generally worth scaling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.054019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.054019Z digest=sha256:55fb35bb200154af813e852c8713f96f0764cae90819d0d6274edba36cfe7bba

Observation f0bc3db2-a78a-4178-b27e-a952fae1a53a · outbound

This paper cites Efficient reasoning for large reasoning language models via certainty-guided reflec- tion suppression.arXiv preprint arXiv:2508.05337, 2025.

Membership Inference Attacks on Tokenizers of Large Language Models Efficient reasoning for large reasoning language models via certainty-guided reflec- tion suppression.arXiv preprint arXiv:2508.05337, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.166718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.166718Z digest=sha256:f4a40fc26f888047961d6ac50e5b92f30906b04df281d40fbbd7994e22f312f2

Observation 8eb61202-b734-4db2-9f5d-171715e968be · outbound

This paper cites Tokenizer.

Membership Inference Attacks on Tokenizers of Large Language Models Tokenizer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.262904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.262904Z digest=sha256:86303489cdf3e07d34ac21b826b598aa00e89cbee476e2046f9d4078bf56341f

Observation e2893360-396c-4497-a880-4c6c1d9767cc · outbound

This paper cites Codeparrot github code dataset.

Membership Inference Attacks on Tokenizers of Large Language Models Codeparrot github code dataset

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.329421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.329421Z digest=sha256:716ff63e23b3d075fa4d3e81e0e05fba086eb6deafe8f8fa8c636fa5df748a3a

Observation d1b47ce6-aa7c-4ca6-a79c-7922e1016dcf · outbound

This paper cites Practical blind membership inference attack via differential compar- isons.

Membership Inference Attacks on Tokenizers of Large Language Models Practical blind membership inference attack via differential compar- isons

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.413438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.413438Z digest=sha256:5b10b79c9b447f1417e3578730b62173af317b6672f8aa387504f58ada61aff1

Observation a3ed5296-d8e7-4df0-8f14-c6f16d8cddc0 · outbound

This paper cites Memguard: Defend- ing against black-box membership inference attacks via adversarial examples.

Membership Inference Attacks on Tokenizers of Large Language Models Memguard: Defend- ing against black-box membership inference attacks via adversarial examples

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.501799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.501799Z digest=sha256:e3e0579710d26092e67c7ee5be8097992fb65f1cdaf7d86ed44c19e304f8e3b7

Observation 8af5eba4-a178-440e-8c0e-38cef2378ecb · outbound

This paper cites Scaling Laws for Neural Language Models.

Membership Inference Attacks on Tokenizers of Large Language Models Scaling Laws for Neural Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.608484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.608484Z digest=sha256:ac339832a6cea390de2ffac005a8897113aa1acef26faeba9d07df710212f723

Observation 279d3227-3f6e-4198-8176-d1b049242167 · outbound

This paper cites Stolen memories: Leveraging model memorization for calibrated{White- Box} membership inference.

Membership Inference Attacks on Tokenizers of Large Language Models Stolen memories: Leveraging model memorization for calibrated{White- Box} membership inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.700474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.700474Z digest=sha256:8ba3309073c37458eb6c8f6f5b52dc163e7f2ec27a8f3246ef3dd2307d848d51

Observation 76a9f61f-abd6-44ab-9944-6bcc8ebd7a64 · outbound

This paper cites Se- qmia: sequential-metric based membership inference attack.

Membership Inference Attacks on Tokenizers of Large Language Models Se- qmia: sequential-metric based membership inference attack

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.769015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.769015Z digest=sha256:09ac95fd0cd690d4cc3bf69e785ced0d07853ba563bffe863b08684f0deeee4f

Observation 5ef775ed-58c5-498f-9a6d-baa0e9342e2c · outbound

This paper cites Enhanced label-only membership inference attacks with fewer queries.

Membership Inference Attacks on Tokenizers of Large Language Models Enhanced label-only membership inference attacks with fewer queries

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:15.884237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:15.884237Z digest=sha256:020571d08c84b426e0f976e9436e9057b29692409dcf0756f44a26527a367f19

Observation 21eed840-cee6-4aec-8338-bdc8e24887a4 · outbound

This paper cites Mem- bership inference attacks and defenses in classification models.

Membership Inference Attacks on Tokenizers of Large Language Models Mem- bership inference attacks and defenses in classification models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.022176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.022176Z digest=sha256:5e1e6613150f13b1a3d04d0e00a411c905695e6bc87f0c3952db0a5f02b358c6

Observation bdb8a7d8-0346-4082-ab60-e33047b1b90b · outbound

This paper cites Large Language Models Can Be Strong Differentially Private Learners.

Membership Inference Attacks on Tokenizers of Large Language Models Large Language Models Can Be Strong Differentially Private Learners

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.101013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.101013Z digest=sha256:f0ba7baabd7154b3d6b1d870c710ba34163a01bf5bf177271ae15b877db3801e

Observation 6c7a5148-7e2e-4d9b-bf33-a21a6259c3de · outbound

This paper cites Membership leakage in label-only exposures.

Membership Inference Attacks on Tokenizers of Large Language Models Membership leakage in label-only exposures

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.184313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.184313Z digest=sha256:a5c17288974e0831cb6050833cd0d32cede6c0b40f9e23501d2df5578eacf814

Observation d565d62b-72d0-436d-b87f-6575bdc965fa · outbound

This paper cites SuperBPE: Space travel for language models.

Membership Inference Attacks on Tokenizers of Large Language Models SuperBPE: Space travel for language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.275636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.275636Z digest=sha256:3888d2fe70305f2bc3671a17580645b39ed090a4ef908e367dd6bbe94da54856

Observation 26dfcea7-eefe-4573-b0ca-10a9ca6486cd · outbound

This paper cites Please tell me more: Privacy impact of explainability through the lens of membership inference attack.

Membership Inference Attacks on Tokenizers of Large Language Models Please tell me more: Privacy impact of explainability through the lens of membership inference attack

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.360687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.360687Z digest=sha256:df01548e6d1bed0e6efe86114d9749926b073184af74907787b34a861623b381

Observation 1cf32503-e51d-48dc-b149-1ca34bcbb4d4 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Membership Inference Attacks on Tokenizers of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.460478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.460478Z digest=sha256:63f9ab1ddbf4f0067d4f78cc9d6be9679bc838589c009295796666aed31ebb6e

Observation 8eeb39b4-8961-4b75-84b6-4d2cfb4f6bb9 · outbound

This paper cites Visu- alizing data using t-sne.Journal of machine learning research, 9(Nov):2579–2605, 2008.

Membership Inference Attacks on Tokenizers of Large Language Models Visu- alizing data using t-sne.Journal of machine learning research, 9(Nov):2579–2605, 2008

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.616036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.616036Z digest=sha256:881b803bdcb00af28f6f41db8030edde07a439ed2895c1f68125d6e057bfc081

Observation a12af4f5-b25c-41ab-b928-be75d08e663c · outbound

This paper cites Llm dataset inference: Did you train on my dataset?Advances in Neural Information Pro- cessing Systems, 37:124069–124092, 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Llm dataset inference: Did you train on my dataset?Advances in Neural Information Pro- cessing Systems, 37:124069–124092, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.748379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.748379Z digest=sha256:d5a33546315c0c63e0ebc6fe12f898550a4b0853cce4dfab1b1d3ec15a6e7c56

Observation 4e79cc26-6e8e-4630-918c-6d305bfe0198 · outbound

This paper cites Dataset Inference: Ownership Resolution in Machine Learning.

Membership Inference Attacks on Tokenizers of Large Language Models Dataset Inference: Ownership Resolution in Machine Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.938718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.938718Z digest=sha256:f677331504e97fdb6526aea857d5e2895e87a3954ce237399e8b960039503354

Observation c65b1c6d-c3e5-49ad-b80d-604c169f334a · outbound

This paper cites Tokens used by gpt-4 probably come from the reddit users.

Membership Inference Attacks on Tokenizers of Large Language Models Tokens used by gpt-4 probably come from the reddit users

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.071614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.071614Z digest=sha256:e7c98f6cbf4f3defc2eda95ead26b406710f92580675aa8400084442f69888bd

Observation a66197b2-16c6-4989-9f59-589f0b43630b · outbound

This paper cites LLMs on the line: Data determines loss-to-loss scaling laws.

Membership Inference Attacks on Tokenizers of Large Language Models LLMs on the line: Data determines loss-to-loss scaling laws

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.199788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.199788Z digest=sha256:816f6a2ef235644727bee5f81ae6f83509a81e7b93bce47d45e2cfac1ee10016

Observation b3521f76-742b-4828-9233-99b2a9e4ada4 · outbound

This paper cites Did the neurons read your book? document-level membership inference for large language models.

Membership Inference Attacks on Tokenizers of Large Language Models Did the neurons read your book? document-level membership inference for large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.334128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.334128Z digest=sha256:0385e78b9432fa59920815f7e92be22758bb41d8f7a40a335ba1c8d2b24c4d3d

Observation b912c77d-5959-4f74-bc61-7b783e172bce · outbound

This paper cites Sok: Membership inference attacks on llms are rushing nowhere (and how to fix it).

Membership Inference Attacks on Tokenizers of Large Language Models Sok: Membership inference attacks on llms are rushing nowhere (and how to fix it)

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.463752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.463752Z digest=sha256:e65ea3e3646077b5e8ee88f23abdba6f0667b42d567931b8ec8ad7942ed3a3b2

Observation c56e5eb5-d70d-4d4f-bd27-e1fae8896bd6 · outbound

This paper cites Pointer Sentinel Mixture Models.

Membership Inference Attacks on Tokenizers of Large Language Models Pointer Sentinel Mixture Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.649968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.649968Z digest=sha256:7ff2c107c9e256c05d239dd3afa0f60d5885b9858ebc06cd9738c488d3a93467

Observation eb9b0c07-cbf7-4335-bfb8-90354e5c857d · outbound

This paper cites Large Language Models: A Survey.

Membership Inference Attacks on Tokenizers of Large Language Models Large Language Models: A Survey

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.742670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.742670Z digest=sha256:aeb7d09cc3aeda524f3259bb73f431555eb34429e8ed8898f0849843fc4d27f5

Observation 92dacadc-918c-4b91-b100-2903c2973956 · outbound

This paper cites Scaling data-constrained language models.Advances in Neural Information Processing Systems, 36:50358– 50376, 2023.

Membership Inference Attacks on Tokenizers of Large Language Models Scaling data-constrained language models.Advances in Neural Information Processing Systems, 36:50358– 50376, 2023

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.811670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.811670Z digest=sha256:565a08b9f09b99b4969cf16a99221b5166b24d3c872a87e934d59b445cd5d6d5

Observation 62cbb5c0-7bcd-4317-855d-638d16060ac2 · outbound

This paper cites Com- prehensive privacy analysis of deep learning: Passive and active white-box inference attacks against central- ized and federated learning.

Membership Inference Attacks on Tokenizers of Large Language Models Com- prehensive privacy analysis of deep learning: Passive and active white-box inference attacks against central- ized and federated learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.912650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.912650Z digest=sha256:c0c0583a12ee90bab9e9f5ee75e36f33a74a98b3a6f0efbbbd92371fc546b28c

Observation c845895b-7816-4174-be76-64db55a21a2a · outbound

This paper cites Reddit sues anthropic over its data scraping to train large language models.

Membership Inference Attacks on Tokenizers of Large Language Models Reddit sues anthropic over its data scraping to train large language models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.066530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.066530Z digest=sha256:402edec8860dfdefcd15093e520eac74237ae56ab7d73ad4bbf976a06561fbf8

Observation 0b8b1236-ce26-4eb9-8e53-6638b6b9ebf8 · outbound

This paper cites System card of chatgpt-o1.

Membership Inference Attacks on Tokenizers of Large Language Models System card of chatgpt-o1

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.205641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.205641Z digest=sha256:b6fff74854f5bc3a4384969b4ecf3cd56d563cb808f5fce1560e47dce7470f33

Observation 1fa7d70f-00bd-4771-972d-3c31d89bf895 · outbound

This paper cites tiktoken: Tokenizer for openai models.

Membership Inference Attacks on Tokenizers of Large Language Models tiktoken: Tokenizer for openai models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.332969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.332969Z digest=sha256:d1d134622a6a29fd2db019c079e16a685cf131ee3f308b60f2ea90e234ad6b4c

Observation cff1adc0-6037-44a8-8ea4-f8c1590afbdd · outbound

This paper cites Black-box Membership Inference Attacks against Fine-tuned Diffusion Models.

Membership Inference Attacks on Tokenizers of Large Language Models Black-box Membership Inference Attacks against Fine-tuned Diffusion Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.538099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.538099Z digest=sha256:aece9418b7617caeb6138dca7afbd4a175c10e73c3faa7c3f58e94c4f37648b4

Observation 309cdaaa-72d0-4beb-bf5c-35641c9f9867 · outbound

This paper cites White-box Membership Inference Attacks against Diffusion Models.

Membership Inference Attacks on Tokenizers of Large Language Models White-box Membership Inference Attacks against Diffusion Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.669175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.669175Z digest=sha256:98986410de4ba3a06842beaa74826b327831f608c65f59aa9ea6053bd5eecf3f

Observation e1ce9bac-38e7-45ed-8447-d787aea53c2c · outbound

This paper cites Zipf’s word frequency law in natural language: A critical review and future direc- tions.Psychonomic bulletin & review, 21(5):1112– 1130, 2014.

Membership Inference Attacks on Tokenizers of Large Language Models Zipf’s word frequency law in natural language: A critical review and future direc- tions.Psychonomic bulletin & review, 21(5):1112– 1130, 2014

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.847087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.847087Z digest=sha256:8b09d85fc613a4db5bec9be8f4de6f5868251c631c89fe9f20b25e02e7c25a3c

Observation 49d85827-7f33-4635-8472-ae11687501c3 · outbound

This paper cites Scaling up membership inference: When and how attacks succeed on large language mod- els.

Membership Inference Attacks on Tokenizers of Large Language Models Scaling up membership inference: When and how attacks succeed on large language mod- els

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.013701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.013701Z digest=sha256:558b2eda249a960d6200564f41a7c5c428b9431c7f33dea84cdce1c9b3f7e6de

Observation 5fde49ab-bb2d-4153-96ea-f02388250892 · outbound

This paper cites Text mining: use of tf-idf to examine the relevance of words to docu- ments.International journal of computer applications, 181(1):25–29, 2018.

Membership Inference Attacks on Tokenizers of Large Language Models Text mining: use of tf-idf to examine the relevance of words to docu- ments.International journal of computer applications, 181(1):25–29, 2018

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.137368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.137368Z digest=sha256:d444171ec017d8d2dea9fc8ac6eb47115cfefb9da32710cb38709cf4c6eba1b9

Observation 5d5a780b-0dd7-447e-9389-f11170a780ec · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Membership Inference Attacks on Tokenizers of Large Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.262666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.262666Z digest=sha256:3dab8b0a401b58fa0debb69a19ba5ae65bf7ec4b3b9f67da5231b4971c4eec5d

Observation e33d4ce0-8b00-4cf1-b85d-cf3cf1ecf926 · outbound

This paper cites Using tf-idf to determine word rele- vance in document queries.

Membership Inference Attacks on Tokenizers of Large Language Models Using tf-idf to determine word rele- vance in document queries

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.347720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.347720Z digest=sha256:6b54871c331872d4c6b85594d88130f6b325767749b59ef5ff1bed04fad822ab

Observation d7523255-b8bc-45f6-acb1-d4454dba0d7d · outbound

This paper cites Gpqa: A graduate- level google-proof q&a benchmark.

Membership Inference Attacks on Tokenizers of Large Language Models Gpqa: A graduate- level google-proof q&a benchmark

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.498048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.498048Z digest=sha256:a2385d304e64a2e7dfc09abbb2bcd6734a7931179c6939770648d1db722d6923

Observation da2ec6eb-684f-4777-b0de-5350c18d96d6 · outbound

This paper cites Self- comparison for dataset-level membership inference in large (vision-) language model.

Membership Inference Attacks on Tokenizers of Large Language Models Self- comparison for dataset-level membership inference in large (vision-) language model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.674528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.674528Z digest=sha256:a51db09f9db916cfd68475643ea648093e9ee76495c3c6e998aee018dfdcba21

Observation 1fa1e540-20f2-4ee0-99e0-f8fcde0e78bb · outbound

This paper cites Learning representations by back- propagating errors.nature, 323(6088):533–536, 1986.

Membership Inference Attacks on Tokenizers of Large Language Models Learning representations by back- propagating errors.nature, 323(6088):533–536, 1986

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.797954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.797954Z digest=sha256:707201a48ca490a4d01c26da169198c985f4384df99cbe104d496f1ba90f04fb

Observation bd4cda93-0789-409a-9ae8-c4b084d0833e · outbound

This paper cites White-box vs black-box: Bayes optimal strategies for membership inference.

Membership Inference Attacks on Tokenizers of Large Language Models White-box vs black-box: Bayes optimal strategies for membership inference

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.969845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.969845Z digest=sha256:e79e004e192b7c04c7c80e1cfe7dff413c09ef78952cf331d63b5a2b0d18a777

Observation b145e816-191e-42a5-8bd6-67fd13cedbb0 · outbound

This paper cites Simple and effec- tive masked diffusion language models.Advances in Neural Information Processing Systems, 37:130136– 130184, 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Simple and effec- tive masked diffusion language models.Advances in Neural Information Processing Systems, 37:130136– 130184, 2024

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.139762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.139762Z digest=sha256:02dc142168806ec6fe2dc075b23a917df13ba10cefe4fa5de308c8341cfc2843

Observation 86ed5bd1-17ee-4c7d-a0ec-91eca9e799fa · outbound

This paper cites ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models.

Membership Inference Attacks on Tokenizers of Large Language Models ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.294107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.294107Z digest=sha256:eb432809b6198330cdd3f7a8babc7b3de5b8c8e7e89bb282ea4c0e372dae082e

Observation 974bae74-aa3a-4128-a61b-b6c70b93111f · outbound

This paper cites Genomic privacy and limits of individual detection in a pool.Nature genetics, 41(9):965–967, 2009.

Membership Inference Attacks on Tokenizers of Large Language Models Genomic privacy and limits of individual detection in a pool.Nature genetics, 41(9):965–967, 2009

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.389642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.389642Z digest=sha256:cf494858f838c711ef841527921d551ba471301be31f2945d5b80f9c2359d633

Observation 00bc5630-9f7c-4a85-a05f-ed85a5ace7e8 · outbound

This paper cites Language models are greedy reasoners: A systematic formal analysis of chain-of-thought.

Membership Inference Attacks on Tokenizers of Large Language Models Language models are greedy reasoners: A systematic formal analysis of chain-of-thought

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.478792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.478792Z digest=sha256:e275c3167cadf3a3b924aad4f1b36ab0dcb5154173f7fc22c0f3ebd854b2d42c

Observation e0710dc0-34b3-43f1-84b8-160776b7f55a · outbound

This paper cites Spearman’s rank correlation coeffi- cient.Bmj, 349, 2014.

Membership Inference Attacks on Tokenizers of Large Language Models Spearman’s rank correlation coeffi- cient.Bmj, 349, 2014

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.596848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.596848Z digest=sha256:a6d2d1156e7920b608f033a88dd85e70481bf1916f98d10abd23b1a8f718dcc3

Observation 5f2526e0-ee78-4dd7-bf73-352210dfe31c · outbound

This paper cites Detecting pretraining data from large language models.

Membership Inference Attacks on Tokenizers of Large Language Models Detecting pretraining data from large language models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.679669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.679669Z digest=sha256:38b22187e87bf59e680b6b6a47cf254159e7016498c97403ca05981583794ec1

Observation 99600e56-1ee2-4f41-9361-43aa1acf2b31 · outbound

This paper cites Byte pair encoding: A text compression scheme that accelerates pattern matching.

Membership Inference Attacks on Tokenizers of Large Language Models Byte pair encoding: A text compression scheme that accelerates pattern matching

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.777679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.777679Z digest=sha256:09d4596ba078184e99d2a2444b260396db7f76e3dd456ccf547e3a872e682632

Observation 8044ccf4-3fbf-40a6-8cac-dc6d50822373 · outbound

This paper cites Membership inference attacks against machine learning models.

Membership Inference Attacks on Tokenizers of Large Language Models Membership inference attacks against machine learning models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.858316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.858316Z digest=sha256:66b4167b5db658d631ae2453cc88749a517243ea868b9a88fd616cd9bb04f8dd

Observation d5b7db1b-baf2-4bc6-a760-310ad373ad0c · outbound

This paper cites Spacebyte: Towards deleting tokeniza- tion from large language modeling.Advances in Neural Information Processing Systems, 37:124925– 124950, 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Spacebyte: Towards deleting tokeniza- tion from large language modeling.Advances in Neural Information Processing Systems, 37:124925– 124950, 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.929826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.929826Z digest=sha256:a21efc01af85106b09854200a3647e7001c536b75461ff352101de7c43a9c246

Observation e5676652-d172-4871-93f1-c0deacbfa35f · outbound

This paper cites A statistical interpretation of term specificity and its application in retrieval.Journal of documentation, 28(1):11–21, 1972.

Membership Inference Attacks on Tokenizers of Large Language Models A statistical interpretation of term specificity and its application in retrieval.Journal of documentation, 28(1):11–21, 1972

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.975566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.975566Z digest=sha256:69e91e08ff07e2742e15f6b6345b730757940975818e3819826d9eff6d6f10db

Observation 142225ce-1cc3-42ad-86a2-28a7c0f5c872 · outbound

This paper cites Scaling laws with vocabulary: Larger mod- els deserve larger vocabularies.Advances in Neural Information Processing Systems, 37:114147–114179, 2024.

Membership Inference Attacks on Tokenizers of Large Language Models Scaling laws with vocabulary: Larger mod- els deserve larger vocabularies.Advances in Neural Information Processing Systems, 37:114147–114179, 2024

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:21.056858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:21.056858Z digest=sha256:10f358c39258b2c37cea7d8f56ee84b8094ed1a7a29ad918532a12c2dbe8afff

Observation 4d7019a2-f1f9-4993-8c01-4678c4699916 · outbound

This paper cites On the vulnerability of text sanitization.

Membership Inference Attacks on Tokenizers of Large Language Models On the vulnerability of text sanitization

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:21.198886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:21.198886Z digest=sha256:3e042d6a231e8199e038a6c20e50b1cbdd10864ad4c1776b34c544fc5840bf18

Observation 593d9036-5fb9-4c61-8794-aa3d063db7f2 · outbound

This paper cites Inferdpt: Privacy-preserving inference for black-box large language models.IEEE Transactions on Dependable and Secure Computing, 2025.

Membership Inference Attacks on Tokenizers of Large Language Models Inferdpt: Privacy-preserving inference for black-box large language models.IEEE Transactions on Dependable and Secure Computing, 2025

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:21.374299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:21.374299Z digest=sha256:dd9e27ab2a93b8cb94704a049923e84cef8181edf371f99cb4a9b8919a216d84

Observation 027c7336-e9db-4dfa-998f-53f8d225a299 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Membership Inference Attacks on Tokenizers of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:21.490138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:21.490138Z digest=sha256:ad4b3fa61540255cd1eb4e053fe73cb887924cd71e0b25b020b2d0a569fdef7a

Pith citing papers

Observation e572a211-6c16-435b-9491-e5c3ebc5d1ff · inbound

Security Considerations for Multi-agent Systems cites this paper.

Security Considerations for Multi-agent Systems Membership Inference Attacks on Tokenizers of Large Language Models

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.092231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T14:12:14.160789Z digest=sha256:2586cdb9f3cd044b6d0a91bbd57545a75417f058ac7b26f2ae3f679c8468b2c5