Pith. sign in

Paper Citation Record · LEDGER

Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2407.16607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.16607 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:14:19.151913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T08:07:34.333897Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fbace07d-c268-44ca-8942-4164bdf7ec8f · inbound

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs cites this paper.

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:31.759038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:51:31.759038Z digest=sha256:edab5845af0a5bb4ddd625d70b9b7f11e07a0f6276a1b6df0f304bd080b18a43

Observation baf94fca-7108-4f29-bf35-a136d0a494e4 · inbound

Cross-Lingual Transfer of Debiasing and Detoxification in Multilingual LLMs: An Extensive Investigation cites this paper.

Cross-Lingual Transfer of Debiasing and Detoxification in Multilingual LLMs: An Extensive Investigation Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:35:40.905103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:35:40.905103Z digest=sha256:4eea719d28476900793a6689be306bde5a80cef3cd446f9fb3d4e1eeb25b7cd0

Observation aea6de6c-d97c-442a-ac0e-1e1a656a6954 · inbound

Learning Dynamics in Continual Pre-Training for Large Language Models cites this paper.

Learning Dynamics in Continual Pre-Training for Large Language Models Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.151913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.151913Z digest=sha256:6b79f969ff406d72e89be4793f6be1827aef01a70f4eee08065f5ff5fad28b2c

Observation 9df92565-2a41-47f7-bcb1-195a6c33393d · inbound

TokAlign: Efficient Vocabulary Adaptation via Token Alignment cites this paper.

TokAlign: Efficient Vocabulary Adaptation via Token Alignment Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:27.856130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:06:27.856130Z digest=sha256:f0a5420a609f53225b6c982133149d09ca058aaa40bec14960c97085cc9cff9b

Observation 8eb38931-7a71-4cb8-bc87-e83f0f58b84f · inbound

When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity cites this paper.

When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:09:22.668034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:09:22.668034Z digest=sha256:b74f215d39defb402d33e402d58e481b8ad4e3943d467c074e1b5557eeea588d

Observation 59aeb56a-efc9-4be8-9748-245687544735 · inbound

Proxy Compression for Language Modeling cites this paper.

Proxy Compression for Language Modeling Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:07:34.335710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:03:01.274736Z digest=sha256:d52a6bf5d86b8cb49f8e5b13fec63d6529f8a63a8bbef92d1bee54a397bfd199

Observation ab158432-9aa0-4370-9329-649ec41d4444 · inbound

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment cites this paper.

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:37:52.289945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T19:35:10.026293Z digest=sha256:672120c134377b4525830de959a25908fafeb57307d5500b378eeb57aacda6bc