Pith. sign in

Paper Citation Record · LEDGER

Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.05417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05417 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:47.767570Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T08:02:23.834546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation adaaa3ea-031f-43b5-89cb-959510ce716f · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.842520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:c41547c6df31fefe825be4f2555ee1264a132bac6e04fe3d9ff865210be0ebb3

Observation ecc8165b-820d-4d01-a388-e17ce8870faf · inbound

Bit-level BPE: Below the byte boundary cites this paper.

Bit-level BPE: Below the byte boundary Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:47.767570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:47.767570Z digest=sha256:fb67fd70163df82efe5720b176d7fd7e9c9bfe31492d896bcab88aadcc6e9130

Observation 8bdd52c9-3500-4ace-b467-35cb864fa7e0 · inbound

FLEXITOKENS: Flexible Tokenization for Evolving Language Models cites this paper.

FLEXITOKENS: Flexible Tokenization for Evolving Language Models Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:05.175224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T05:10:19.093601Z digest=sha256:12c77f1dbfb1eb49e3b3b2e419b5649877ad907de7d2c94463755a7039f9e60c

Observation 486d7ce7-cc3d-44ce-a922-d19eba8f54a5 · inbound

Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models cites this paper.

Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:43:22.885519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:43:22.885519Z digest=sha256:8bfac23bf5a49ff4ac9e178110a2fdc8c98ea29f0fe29e51df4d7f570b179012

Observation 7c487d96-0e6d-424d-af09-9c94b48633a7 · inbound

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens cites this paper.

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:51:25.782408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T13:33:21.840958Z digest=sha256:3c7735f65f259b0bc72ad850947a0ec2d34734940442c6ebd11f3ae275852947

Observation 7861d0da-8c14-4247-a758-834f7c2ff845 · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.075646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:5ed0e43bfa522a36f3e707cc904d9715b633e4a280213a39dd82c4090868c0db